Definition
Voice AI Pipeline
A voice agent chains three models: speech-to-text transcribes the caller, an LLM reasons and drafts a reply, and text-to-speech renders it as audio. Each stage adds latency, and human conversation expects a reply within a few hundred milliseconds, so every stage must stream into the next. Around 700 milliseconds to first audio feels like a natural pause; below 500 feels real-time. Silence during tool calls is deadly.

Explained in
Chapter 19: Voice and Music AI
The AI revolution isn't just text and images. It's sound.
Related terms