面向智能体的实时语音AI技术栈:架构指南

Deepgram Learn ·

一份技术架构指南,阐述生产级语音智能体对实时语音AI的延迟要求,援引影响交互质量感知的延迟阈值、端点检测设计中的权衡,并基于用户研究与真实通话者反馈设定服务等级目标(SLO)。 阅读 3 条观点,查看支持证据与原始来源。

理解这篇

3 个要点

综合解读

  1. 一秒转述时限

    生产级语音智能体必须在用户话语结束约一秒钟内开始应答,否则该轮对话将显得中断或失效。

    支持这项说法 1

    它约有一秒钟时间启动应答,否则该轮对话将显得中断或失效。

    Jose Nicholas Francisco · 段落 1

    原始摘录
    It has about a second to start talking before the exchange feels broken.
    上下文

    生产级语音智能体将传输、语音识别、语言模型、语音合成及编排等环节串联为一次实时交互,各层共享同一时钟。

    原始上下文

    A production voice agent chains transport, speech recognition, a language model, speech synthesis, and orchestration into a single real-time exchange, and each layer draws from the same shared clock.

    回到原文语境 →
  2. 600–800毫秒:用户继续交互意愿明显下降区间

    用户继续交互的意愿在600毫秒后即开始下降,在700–800毫秒区间内显著降低。

    支持这项说法 1

    用户继续交互的意愿在600毫秒后即开始下降,在700–800毫秒区间内明显降低。

    Jose Nicholas Francisco · 段落 2

    原始摘录
    Perceived willingness begins to drop after 600 ms and steps down significantly from 700 to 800 ms.
    上下文

    传输、语音识别、端点检测、语言模型(LLM)输出首个token、文本转语音(TTS)返回首个字节,以及网络回传跳转——每个环节均占用该延迟预算的一部分。本指南将逐层拆解各环节开销,并指出可优化的具体环节,以确保语音AI技术栈整体满足该预算约束。

    原始上下文

    Transport, recognition, endpointing, the LLM's first token, TTS first byte, and the return network hop each take a slice of that budget. This guide breaks down what each layer costs and where you can cut to keep your voice AI stack inside it.

    回到原文语境 →
  3. P95 低于 1,000 毫秒的目标

    生产环境中单轮延迟的可行目标是 P95 低于约 1,000 毫秒;超过几秒后,通话者便开始失去参与感;应采用基于百分位数的 SLO,而非平均值,并依据用户研究和真实通话者反馈进行校准。

    支持这项说法 1

    将第95百分位(P95)响应延迟目标设为约1000毫秒;超过数秒的延迟,用户通常开始失去参与感。应基于百分位数(而非平均值)设定服务水平目标(SLO);校准可参考IEEE关于语音界面中用户感知响应延迟的研究,再结合自身用户的实际反馈验证和确认具体阈值。

    Jose Nicholas Francisco · 段落 97

    原始摘录
    Target a P95 turn latency under about 1,000 ms, and treat anything beyond a couple of seconds as the point where callers start to disengage. Set percentile SLOs, not averages, and calibrate against user-study research like this IEEE study on perceived response delay in voice interfaces, then confirm the specific thresholds against your own caller feedback.
    回到原文语境 →

语音 Agent:什么时候可以回应?

编辑导读 · 依据另引的官方文档 ·

分别处理转写片段定稿、检测到停顿和正式结束回合。提前准备的回复不等于已经可以播出。

Nova:is_final 与 speech_final 有什么区别?

is_final 表示一段音频的转写定稿;speech_final 表示检测到停顿。按顺序累积每个 final 片段,包括伴随 speech_final 的最后一段,再消费并清空话语缓冲区。不要只保留最后一段,也不要把仍会变化的 interim 文本作为另一段 final 追加。停顿不证明用户已说完自己的意思。

Deepgram · Nova endpointing and interim results ↗

Flux:EagerEndOfTurn → TurnResumed → EndOfTurn

配置 eager_eot_threshold 后,EagerEndOfTurn 可用于开始准备回复。TurnResumed 时取消并作废旧草稿,等待新的 eager 或 EndOfTurn 事件;只播出当前已结束回合的回复。eot_timeout_ms 可在静音超时后强制触发 EndOfTurn。

Deepgram · Flux eager end of turn ↗ Deepgram · Flux end-of-turn configuration ↗

取消后,防止旧草稿进入当前回复

实现建议:把生成与播放绑定到当前回合及草稿版本。取消后,即使旧请求未能停止,也要拒绝迟到的旧版本结果。不要因为生成完成就播放提前准备的草稿。

Deepgram · Flux eager end of turn ↗

合成事件验收草案

  • 多个 Nova final:按顺序保留片段,不丢失、不重复。
  • Flux eager 后继续说话或改口:作废首稿,采用修订后的回合。
  • 正常 EndOfTurn:当前回复只播出一次。
  • 超时触发 EndOfTurn:走相同结束流程,不重复播出。
  • 取消后旧结果迟到:进入播放前丢弃。

这些合成事件测试尚未执行。它们是验收建议,不是经过实测的语音 Agent 实现。

关键段落3

带明确归属与语境的原文片段。打开原始文本核查出处。

生产环境延迟服务等级目标(SLO)

P95 低于 1,000 毫秒的目标

将第95百分位(P95)响应延迟目标设为约1000毫秒;超过数秒的延迟,用户通常开始失去参与感。应基于百分位数(而非平均值)设定服务水平目标(SLO);校准可参考IEEE关于语音界面中用户感知响应延迟的研究,再结合自身用户的实际反馈验证和确认具体阈值。

原始摘录
Target a P95 turn latency under about 1,000 ms, and treat anything beyond a couple of seconds as the point where callers start to disengage. Set percentile SLOs, not averages, and calibrate against user-study research like this IEEE study on perceived response delay in voice interfaces, then confirm the specific thresholds against your own caller feedback.
实时语音AI延迟预算

一秒转述时限

它约有一秒钟时间启动应答,否则该轮对话将显得中断或失效。

原始摘录
It has about a second to start talking before the exchange feels broken.
上下文

生产级语音智能体将传输、语音识别、语言模型、语音合成及编排等环节串联为一次实时交互,各层共享同一时钟。

原始上下文

A production voice agent chains transport, speech recognition, a language model, speech synthesis, and orchestration into a single real-time exchange, and each layer draws from the same shared clock.

感知延迟阈值

600–800毫秒:用户继续交互意愿明显下降区间

用户继续交互的意愿在600毫秒后即开始下降,在700–800毫秒区间内明显降低。

原始摘录
Perceived willingness begins to drop after 600 ms and steps down significantly from 700 to 800 ms.
上下文

传输、语音识别、端点检测、语言模型(LLM)输出首个token、文本转语音(TTS)返回首个字节,以及网络回传跳转——每个环节均占用该延迟预算的一部分。本指南将逐层拆解各环节开销,并指出可优化的具体环节,以确保语音AI技术栈整体满足该预算约束。

原始上下文

Transport, recognition, endpointing, the LLM's first token, TTS first byte, and the return network hop each take a slice of that budget. This guide breaks down what each layer costs and where you can cut to keep your voice AI stack inside it.

来源与研究方法

这些观点均关联原始来源。转述已明确标注,不作为逐字原话展示。

打开转录或来源材料 (在新标签页中打开)报告问题

继续了解这些人物的观点