OPINIONES PÚBLICAS

Jose Nicholas Francisco

2 fuentes · 7 opiniones · 7 temas

Contenido actualizado:

Jose Nicholas Francisco sobre function selection accuracy, latency tolerance in voice interactions, perceived latency thresholds. Explora 7 puntos de vista por tema, con evidencias de 2 fuentes.

Explorar conexiones

Perspectivas por tema

Opiniones atribuidas, ordenadas por fecha de publicación de la fuente. Una muestra de estas intervenciones, no una definición completa de las creencias de la persona.

Las traducciones son para facilitar la lectura; los extractos originales siguen siendo la source evidence.

real-time voice AI latency budget

Ver este tema

One-second turn deadline

A production voice agent must begin speaking within about one second of the user's utterance ending, or the exchange feels broken.

Evidencia a favor

Real-Time Voice AI Stack for Agents: Architecture Guide

Extracto original

It has about a second to start talking before the exchange feels broken.
Contexto

A production voice agent chains transport, speech recognition, a language model, speech synthesis, and orchestration into a single real-time exchange, and each layer draws from the same shared clock.

perceived latency thresholds

Ver este tema

600–800 ms perceived willingness drop

Perceived willingness to continue the interaction begins to decline after 600 ms and drops significantly between 700 and 800 ms.

Evidencia a favor

Real-Time Voice AI Stack for Agents: Architecture Guide

Extracto original

Perceived willingness begins to drop after 600 ms and steps down significantly from 700 to 800 ms.
Contexto

Transport, recognition, endpointing, the LLM's first token, TTS first byte, and the return network hop each take a slice of that budget. This guide breaks down what each layer costs and where you can cut to keep your voice AI stack inside it.

production latency SLOs

Ver este tema

P95 under 1,000 ms target

A realistic production per-turn latency target is P95 under approximately 1,000 ms; beyond a couple of seconds, callers begin to disengage, and percentile-based SLOs—not averages—should be used and calibrated against user studies and real caller feedback.

Evidencia a favor

Real-Time Voice AI Stack for Agents: Architecture Guide

Extracto original

Target a P95 turn latency under about 1,000 ms, and treat anything beyond a couple of seconds as the point where callers start to disengage. Set percentile SLOs, not averages, and calibrate against user-study research like this IEEE study on perceived response delay in voice interfaces, then confirm the specific thresholds against your own caller feedback.

voice agent function calling

Ver este tema

Function calling enables direct answers instead of call transfers

Function calling allows a voice agent to directly answer questions like 'Where’s my order?' rather than transferring the caller to a human agent.

Evidencia a favor

Voice Agent Function Calling: Real-Time Data Lookup Guide

Extracto original

Function calling is how a voice agent answers "where's my order?" instead of transferring the caller to someone who can.
Contexto

One order-status lookup on the inbound telephony agent pattern takes two round trips, three function calls, and a spoken line for every way it can fail.

latency tolerance in voice interactions

Ver este tema

Perceived willingness drops significantly after 600–800 ms of silence

Ratings of perceived willingness drop noticeably at 600 ms and become statistically significant between 700 and 800 ms, based on the 2013 gap-tolerance study by Roberts and Francis.

Evidencia a favor

Voice Agent Function Calling: Real-Time Data Lookup Guide

Extracto original

Ratings of perceived willingness drop off noticeably at 600ms in the classic 2013 gap-tolerance study from Roberts and Francis, and the difference turns statistically significant between 700 and 800ms.
Contexto

Dead air is what the caller hears while your handler works, and it costs you fast. InjectAgentMessage is what covers that silence in the reference pattern.

function selection accuracy

Ver este tema

Tool count degrades selection accuracy: 98% at 10 tools → 88% at 100

The HumanMCP study measured a decline in function selection accuracy from 98% with 10 tools to 88% with 100 tools.

Evidencia a favor

Voice Agent Function Calling: Real-Time Data Lookup Guide

Extracto original

The HumanMCP study measured a drop from 98% at 10 tools to 88% at 100.
Contexto

When the toolset grows, selection accuracy falls. Add functions one at a time and re-check selection accuracy after each. One lookup, three failure lines, and a filler function make a complete first release.

session limits in voice agents

Ver este tema

Two-hour session limit applies regardless of in-flight lookups

Voice agent sessions expire after two hours; an in-flight function call does not extend the session, and the system emits MAXIMUM_SESSION_LENGTH_APPROACHING at 1h 55m and MAXIMUM_SESSION_LENGTH_REACHED at 2h.

Evidencia a favor

Voice Agent Function Calling: Real-Time Data Lookup Guide

Extracto original

Yes, at two hours, and an in-flight lookup won't hold the session open. You get a Warning with MAXIMUM_SESSION_LENGTH_APPROACHING at 1 hour 55 minutes, then an Error with MAXIMUM_SESSION_LENGTH_REACHED at the two-hour limit .
Contexto

KeepAlive doesn't extend it, and the docs don't cover a pending FunctionCallRequest at close. Finish or abandon it, then reconnect with agent.context .

Declaraciones por fecha de la fuente2

Las declaraciones se ordenan por fecha de publicación de la fuente original; las diferencias de redacción no demuestran un cambio de postura.