Real-Time Voice AI Stack for Agents: Architecture Guide

Deepgram Learn ·

A technical architecture guide describing real-time voice AI latency requirements for production voice agents, citing thresholds for perceived interaction quality, tradeoffs in endpointing design, and service-level objectives based on user studies and caller feedback. Read 3 viewpoints with supporting evidence and source links.

Understand this piece

3 key points

Synthesis

  1. One-second turn deadline

    A production voice agent must begin speaking within about one second of the user's utterance ending, or the exchange feels broken.

    Supporting evidence 1

    Original excerpt

    It has about a second to start talking before the exchange feels broken.

    Jose Nicholas Francisco · Paragraph 1

    Context

    A production voice agent chains transport, speech recognition, a language model, speech synthesis, and orchestration into a single real-time exchange, and each layer draws from the same shared clock.

    Read in source context →
  2. 600–800 ms perceived willingness drop

    Perceived willingness to continue the interaction begins to decline after 600 ms and drops significantly between 700 and 800 ms.

    Supporting evidence 1

    Original excerpt

    Perceived willingness begins to drop after 600 ms and steps down significantly from 700 to 800 ms.

    Jose Nicholas Francisco · Paragraph 2

    Context

    Transport, recognition, endpointing, the LLM's first token, TTS first byte, and the return network hop each take a slice of that budget. This guide breaks down what each layer costs and where you can cut to keep your voice AI stack inside it.

    Read in source context →
  3. P95 under 1,000 ms target

    A realistic production per-turn latency target is P95 under approximately 1,000 ms; beyond a couple of seconds, callers begin to disengage, and percentile-based SLOs—not averages—should be used and calibrated against user studies and real caller feedback.

    Supporting evidence 1

    Original excerpt

    Target a P95 turn latency under about 1,000 ms, and treat anything beyond a couple of seconds as the point where callers start to disengage. Set percentile SLOs, not averages, and calibrate against user-study research like this IEEE study on perceived response delay in voice interfaces, then confirm the specific thresholds against your own caller feedback.

    Jose Nicholas Francisco · Paragraph 97

    Read in source context →

    Continue exploring

    production latency SLOs →

When can a voice agent respond?

Editorial guide · Based on the official documentation cited below ·

Keep transcript finalization, pause detection and turn completion separate. Preparing a reply early does not mean it is ready to play.

Nova: is_final vs speech_final

is_final marks a finalized audio segment; speech_final marks a detected pause. Accumulate each final segment in order, including the final segment accompanying speech_final, before consuming and clearing the utterance buffer. Do not keep only the last segment or append mutable interim text as another final segment. A pause is not proof that the speaker has finished their thought.

Deepgram · Nova endpointing and interim results ↗

Flux: EagerEndOfTurn → TurnResumed → EndOfTurn

With eager_eot_threshold configured, EagerEndOfTurn can start drafting. TurnResumed cancels and invalidates that draft; wait for a new eager or EndOfTurn event. Deliver only a response for the current completed turn. eot_timeout_ms can force EndOfTurn after silence.

Deepgram · Flux eager end of turn ↗ Deepgram · Flux end-of-turn configuration ↗

Keep a cancelled draft out of the current reply

Implementation suggestion: bind generation and playback to the current turn and draft revision. After cancellation, reject late results from the old revision, even if its request could not be stopped. Do not play a speculative draft merely because generation finished.

Deepgram · Flux eager end of turn ↗

Synthetic-event acceptance draft

  • Multiple Nova final segments: preserve their order without loss or duplicate text.
  • Flux eager then resumed speech: invalidate the first draft and use the revised turn.
  • Normal EndOfTurn: deliver the current response once.
  • Timeout-forced EndOfTurn: follow the same completion path without double delivery.
  • Old result arriving after cancellation: discard it before it reaches playback.

These synthetic-event cases have not been executed. They are acceptance suggestions, not a tested voice-agent implementation.

Key passages3

Attributed passages with the context to verify them. Open the original text to check the source.

production latency SLOs

P95 under 1,000 ms target

Original excerpt

Target a P95 turn latency under about 1,000 ms, and treat anything beyond a couple of seconds as the point where callers start to disengage. Set percentile SLOs, not averages, and calibrate against user-study research like this IEEE study on perceived response delay in voice interfaces, then confirm the specific thresholds against your own caller feedback.
real-time voice AI latency budget

One-second turn deadline

Original excerpt

It has about a second to start talking before the exchange feels broken.
Context

A production voice agent chains transport, speech recognition, a language model, speech synthesis, and orchestration into a single real-time exchange, and each layer draws from the same shared clock.

perceived latency thresholds

600–800 ms perceived willingness drop

Original excerpt

Perceived willingness begins to drop after 600 ms and steps down significantly from 700 to 800 ms.
Context

Transport, recognition, endpointing, the LLM's first token, TTS first byte, and the return network hop each take a slice of that budget. This guide breaks down what each layer costs and where you can cut to keep your voice AI stack inside it.

Source & methodology

These viewpoints are linked to their original sources. Paraphrases are labeled and are not verbatim quotes.

Open transcript or source material (opens in a new tab)Report an issue

Explore these viewpoints by person