PUBLIC VIEWPOINTS

Jose Nicholas Francisco

2 sources · 7 viewpoints · 7 topics

Content updated:

Jose Nicholas Francisco on function selection accuracy, latency tolerance in voice interactions, perceived latency thresholds. Explore 7 viewpoints by topic, with evidence from 2 sources.

Explore connections

Viewpoints by topic

Attributed viewpoints, ordered by source publication date. A snapshot of these conversations, not a definitive statement of someone’s beliefs.

Translations are for reading; original excerpts remain the evidence.

real-time voice AI latency budget

View this topic

One-second turn deadline

A production voice agent must begin speaking within about one second of the user's utterance ending, or the exchange feels broken.

Supporting evidence

Real-Time Voice AI Stack for Agents: Architecture Guide

Original excerpt

It has about a second to start talking before the exchange feels broken.
Context

A production voice agent chains transport, speech recognition, a language model, speech synthesis, and orchestration into a single real-time exchange, and each layer draws from the same shared clock.

perceived latency thresholds

View this topic

600–800 ms perceived willingness drop

Perceived willingness to continue the interaction begins to decline after 600 ms and drops significantly between 700 and 800 ms.

Supporting evidence

Real-Time Voice AI Stack for Agents: Architecture Guide

Original excerpt

Perceived willingness begins to drop after 600 ms and steps down significantly from 700 to 800 ms.
Context

Transport, recognition, endpointing, the LLM's first token, TTS first byte, and the return network hop each take a slice of that budget. This guide breaks down what each layer costs and where you can cut to keep your voice AI stack inside it.

production latency SLOs

View this topic

P95 under 1,000 ms target

A realistic production per-turn latency target is P95 under approximately 1,000 ms; beyond a couple of seconds, callers begin to disengage, and percentile-based SLOs—not averages—should be used and calibrated against user studies and real caller feedback.

Supporting evidence

Real-Time Voice AI Stack for Agents: Architecture Guide

Original excerpt

Target a P95 turn latency under about 1,000 ms, and treat anything beyond a couple of seconds as the point where callers start to disengage. Set percentile SLOs, not averages, and calibrate against user-study research like this IEEE study on perceived response delay in voice interfaces, then confirm the specific thresholds against your own caller feedback.

voice agent function calling

View this topic

Function calling enables direct answers instead of call transfers

Function calling allows a voice agent to directly answer questions like 'Where’s my order?' rather than transferring the caller to a human agent.

Supporting evidence

Voice Agent Function Calling: Real-Time Data Lookup Guide

Original excerpt

Function calling is how a voice agent answers "where's my order?" instead of transferring the caller to someone who can.
Context

One order-status lookup on the inbound telephony agent pattern takes two round trips, three function calls, and a spoken line for every way it can fail.

latency tolerance in voice interactions

View this topic

Perceived willingness drops significantly after 600–800 ms of silence

Ratings of perceived willingness drop noticeably at 600 ms and become statistically significant between 700 and 800 ms, based on the 2013 gap-tolerance study by Roberts and Francis.

Supporting evidence

Voice Agent Function Calling: Real-Time Data Lookup Guide

Original excerpt

Ratings of perceived willingness drop off noticeably at 600ms in the classic 2013 gap-tolerance study from Roberts and Francis, and the difference turns statistically significant between 700 and 800ms.
Context

Dead air is what the caller hears while your handler works, and it costs you fast. InjectAgentMessage is what covers that silence in the reference pattern.

function selection accuracy

View this topic

Tool count degrades selection accuracy: 98% at 10 tools → 88% at 100

The HumanMCP study measured a decline in function selection accuracy from 98% with 10 tools to 88% with 100 tools.

Supporting evidence

Voice Agent Function Calling: Real-Time Data Lookup Guide

Original excerpt

The HumanMCP study measured a drop from 98% at 10 tools to 88% at 100.
Context

When the toolset grows, selection accuracy falls. Add functions one at a time and re-check selection accuracy after each. One lookup, three failure lines, and a filler function make a complete first release.

session limits in voice agents

View this topic

Two-hour session limit applies regardless of in-flight lookups

Voice agent sessions expire after two hours; an in-flight function call does not extend the session, and the system emits MAXIMUM_SESSION_LENGTH_APPROACHING at 1h 55m and MAXIMUM_SESSION_LENGTH_REACHED at 2h.

Supporting evidence

Voice Agent Function Calling: Real-Time Data Lookup Guide

Original excerpt

Yes, at two hours, and an in-flight lookup won't hold the session open. You get a Warning with MAXIMUM_SESSION_LENGTH_APPROACHING at 1 hour 55 minutes, then an Error with MAXIMUM_SESSION_LENGTH_REACHED at the two-hour limit .
Context

KeepAlive doesn't extend it, and the docs don't cover a pending FunctionCallRequest at close. Finish or abandon it, then reconnect with agent.context .

Statements by source date2

Statements are ordered by the original source publication date; differences in wording do not establish a change of position.