公开观点

Jose Nicholas Francisco

2 份资料 · 7 条观点 · 7 个话题

内容更新于:

Jose Nicholas Francisco 关于函数选择准确率、语音交互中的延迟容忍度、感知延迟阈值的观点。 按话题阅读 7 条观点,核对 2 个来源中的证据。

探索知识关联

按话题查看观点

按来源发布日期整理的个人观点,仅反映这些材料中的表达,不代表其全部立场。

译文仅辅助阅读;核查观点请以原始摘录为准。

实时语音AI延迟预算

查看话题

一秒转述时限

生产级语音智能体必须在用户话语结束约一秒钟内开始应答,否则该轮对话将显得中断或失效。

支持这项说法

面向智能体的实时语音AI技术栈:架构指南

它约有一秒钟时间启动应答,否则该轮对话将显得中断或失效。

原始摘录
It has about a second to start talking before the exchange feels broken.
上下文

生产级语音智能体将传输、语音识别、语言模型、语音合成及编排等环节串联为一次实时交互,各层共享同一时钟。

原始上下文

A production voice agent chains transport, speech recognition, a language model, speech synthesis, and orchestration into a single real-time exchange, and each layer draws from the same shared clock.

感知延迟阈值

查看话题
感知延迟阈值

600–800毫秒:用户继续交互意愿明显下降区间

用户继续交互的意愿在600毫秒后即开始下降,在700–800毫秒区间内显著降低。

支持这项说法

面向智能体的实时语音AI技术栈:架构指南

用户继续交互的意愿在600毫秒后即开始下降,在700–800毫秒区间内明显降低。

原始摘录
Perceived willingness begins to drop after 600 ms and steps down significantly from 700 to 800 ms.
上下文

传输、语音识别、端点检测、语言模型(LLM)输出首个token、文本转语音(TTS)返回首个字节,以及网络回传跳转——每个环节均占用该延迟预算的一部分。本指南将逐层拆解各环节开销,并指出可优化的具体环节,以确保语音AI技术栈整体满足该预算约束。

原始上下文

Transport, recognition, endpointing, the LLM's first token, TTS first byte, and the return network hop each take a slice of that budget. This guide breaks down what each layer costs and where you can cut to keep your voice AI stack inside it.

生产环境延迟服务等级目标(SLO)

查看话题

P95 低于 1,000 毫秒的目标

生产环境中单轮延迟的可行目标是 P95 低于约 1,000 毫秒;超过几秒后,通话者便开始失去参与感;应采用基于百分位数的 SLO,而非平均值,并依据用户研究和真实通话者反馈进行校准。

支持这项说法

面向智能体的实时语音AI技术栈:架构指南

将第95百分位(P95)响应延迟目标设为约1000毫秒;超过数秒的延迟,用户通常开始失去参与感。应基于百分位数(而非平均值)设定服务水平目标(SLO);校准可参考IEEE关于语音界面中用户感知响应延迟的研究,再结合自身用户的实际反馈验证和确认具体阈值。

原始摘录
Target a P95 turn latency under about 1,000 ms, and treat anything beyond a couple of seconds as the point where callers start to disengage. Set percentile SLOs, not averages, and calibrate against user-study research like this IEEE study on perceived response delay in voice interfaces, then confirm the specific thresholds against your own caller feedback.

语音代理函数调用

查看话题
语音代理函数调用

函数调用支持直接回答,避免转接电话

函数调用使语音代理能够直接回答‘我的订单在哪里?’之类的问题,而非将呼叫者转接给人类座席。

支持这项说法

语音代理函数调用指南:实时数据查询

函数调用使语音代理能直接回答‘我的订单在哪里?’,而非将呼叫者转接给人工处理该问题。

原始摘录
Function calling is how a voice agent answers "where's my order?" instead of transferring the caller to someone who can.
上下文

入站电话语音代理模式下的一次订单状态查询需经历两次往返、三次函数调用,并针对每种可能的失败路径生成一句语音反馈。

原始上下文

One order-status lookup on the inbound telephony agent pattern takes two round trips, three function calls, and a spoken line for every way it can fail.

语音交互中的延迟容忍度

查看话题

感知意愿在600–800毫秒静默后显著下降

根据罗伯茨(Roberts)与弗朗西斯(Francis)2013年关于‘间隔容忍度’的研究,感知意愿评分在静默达600毫秒时即明显下降,并在700–800毫秒区间达到统计显著性。

支持这项说法

语音代理函数调用指南:实时数据查询

在罗伯茨(Roberts)与弗朗西斯(Francis)2013年开展的经典‘间隔容忍度’研究中,受试者对意愿的主观评分在600毫秒处明显下降,该差异在700–800毫秒区间内达到统计显著性。

原始摘录
Ratings of perceived willingness drop off noticeably at 600ms in the classic 2013 gap-tolerance study from Roberts and Francis, and the difference turns statistically significant between 700 and 800ms.
上下文

‘死寂’指主叫方在您的处理器执行任务期间听到的静音;这种静音会迅速损害用户体验。InjectAgentMessage 是参考模式中用于填补该静音的机制。

原始上下文

Dead air is what the caller hears while your handler works, and it costs you fast. InjectAgentMessage is what covers that silence in the reference pattern.

函数选择准确率

查看话题
函数选择准确率

工具数从10增至100,选择准确率由98%降至88%

HumanMCP研究发现:函数选择准确率随可用工具数量增加而下降——使用10个工具时为98%,增至100个工具时降至88%。

支持这项说法

语音代理函数调用指南:实时数据查询

HumanMCP 研究测得,在 10 个工具时为 98%,到 100 个工具时下降至 88%。

原始摘录
The HumanMCP study measured a drop from 98% at 10 tools to 88% at 100.
上下文

当工具集规模扩大时,选择准确率会下降。应逐个添加函数,并在每次添加后重新检查选择准确率。一个查找、三条失败行和一个填充函数即可构成完整的首次发布版本。

原始上下文

When the toolset grows, selection accuracy falls. Add functions one at a time and re-check selection accuracy after each. One lookup, three failure lines, and a filler function make a complete first release.

语音代理的会话时限

查看话题

两小时会话时限强制生效,无论是否存在进行中的查询

语音代理会话在两小时后过期;进行中的函数调用不会延长会话,系统在 1h 55m 发出 MAXIMUM_SESSION_LENGTH_APPROACHING,在 2h 发出 MAXIMUM_SESSION_LENGTH_REACHED。

支持这项说法

语音代理函数调用指南:实时数据查询

是的,时限为两小时,进行中的查找不会保持会话开启。在 1 小时 55 分钟时,你会收到一条带有 MAXIMUM_SESSION_LENGTH_APPROACHING 的 Warning;在两小时上限时,则会收到一条带有 MAXIMUM_SESSION_LENGTH_REACHED 的 Error。

原始摘录
Yes, at two hours, and an in-flight lookup won't hold the session open. You get a Warning with MAXIMUM_SESSION_LENGTH_APPROACHING at 1 hour 55 minutes, then an Error with MAXIMUM_SESSION_LENGTH_REACHED at the two-hour limit .
上下文

KeepAlive 无法延长该时限,且文档未涵盖关闭时处于待处理状态的 FunctionCallRequest。完成或放弃它,然后通过 agent.context 重新连接。

原始上下文

KeepAlive doesn't extend it, and the docs don't cover a pending FunctionCallRequest at close. Finish or abandon it, then reconnect with agent.context .

按来源日期阅读2

按原始来源的发布日期排序;措辞不同不代表立场发生变化。