Real-Time Voice AI Stack for Agents: Architecture Guide
Deepgram Learn ·
A technical architecture guide describing real-time voice AI latency requirements for production voice agents, citing thresholds for perceived interaction quality, tradeoffs in endpointing design, and service-level objectives based on user studies and caller feedback. Read 3 viewpoints with supporting evidence and source links.
Lines show the reading structure. Select an idea to read its explanation and evidence.
Synthesis
01
One-second turn deadline
A production voice agent must begin speaking within about one second of the user's utterance ending, or the exchange feels broken.
Supporting evidence 1
Original excerpt
It has about a second to start talking before the exchange feels broken.
Jose Nicholas Francisco · Paragraph 1
Real-Time Voice AI Stack for Agents: Architecture Guide · Aug 25, 2026
Context
A production voice agent chains transport, speech recognition, a language model, speech synthesis, and orchestration into a single real-time exchange, and each layer draws from the same shared clock.
Perceived willingness to continue the interaction begins to decline after 600 ms and drops significantly between 700 and 800 ms.
Supporting evidence 1
Original excerpt
Perceived willingness begins to drop after 600 ms and steps down significantly from 700 to 800 ms.
Jose Nicholas Francisco · Paragraph 2
Real-Time Voice AI Stack for Agents: Architecture Guide · Aug 25, 2026
Context
Transport, recognition, endpointing, the LLM's first token, TTS first byte, and the return network hop each take a slice of that budget. This guide breaks down what each layer costs and where you can cut to keep your voice AI stack inside it.
A realistic production per-turn latency target is P95 under approximately 1,000 ms; beyond a couple of seconds, callers begin to disengage, and percentile-based SLOs—not averages—should be used and calibrated against user studies and real caller feedback.
Supporting evidence 1
Original excerpt
Target a P95 turn latency under about 1,000 ms, and treat anything beyond a couple of seconds as the point where callers start to disengage. Set percentile SLOs, not averages, and calibrate against user-study research like this IEEE study on perceived response delay in voice interfaces, then confirm the specific thresholds against your own caller feedback.
Jose Nicholas Francisco · Paragraph 97
Real-Time Voice AI Stack for Agents: Architecture Guide · Aug 25, 2026
Editorial guide · Based on the official documentation cited below ·
Keep transcript finalization, pause detection and turn completion separate. Preparing a reply early does not mean it is ready to play.
Nova: is_final vs speech_final
is_final marks a finalized audio segment; speech_final marks a detected pause. Accumulate each final segment in order, including the final segment accompanying speech_final, before consuming and clearing the utterance buffer. Do not keep only the last segment or append mutable interim text as another final segment. A pause is not proof that the speaker has finished their thought.
With eager_eot_threshold configured, EagerEndOfTurn can start drafting. TurnResumed cancels and invalidates that draft; wait for a new eager or EndOfTurn event. Deliver only a response for the current completed turn. eot_timeout_ms can force EndOfTurn after silence.
Implementation suggestion: bind generation and playback to the current turn and draft revision. After cancellation, reject late results from the old revision, even if its request could not be stopped. Do not play a speculative draft merely because generation finished.
Target a P95 turn latency under about 1,000 ms, and treat anything beyond a couple of seconds as the point where callers start to disengage. Set percentile SLOs, not averages, and calibrate against user-study research like this IEEE study on perceived response delay in voice interfaces, then confirm the specific thresholds against your own caller feedback.
A production voice agent chains transport, speech recognition, a language model, speech synthesis, and orchestration into a single real-time exchange, and each layer draws from the same shared clock.
Transport, recognition, endpointing, the LLM's first token, TTS first byte, and the return network hop each take a slice of that budget. This guide breaks down what each layer costs and where you can cut to keep your voice AI stack inside it.
Source & methodology
These viewpoints are linked to their original sources. Paraphrases are labeled and are not verbatim quotes.
Jose Nicholas Francisco on function selection accuracy, latency tolerance in voice interactions, perceived latency thresholds. Explore 7 viewpoints by topic, with evidence from 2 sources.