OPINIONS EXPRIMÉES EN PUBLIC

Jose Nicholas Francisco

2 sources · 7 points de vue · 7 sujets

Contenu mis à jour:

Jose Nicholas Francisco sur function selection accuracy, latency tolerance in voice interactions, perceived latency thresholds. Explorez 7 points de vue par thème, avec des éléments tirés de 2 sources.

Explorer les liens

Points de vue par sujet

Points de vue attribués, classés par date de publication de la source. Un aperçu de ces échanges, sans prétendre définir toutes les convictions de la personne.

Les traductions sont destinées à la lecture ; les extraits originaux restent la source evidence.

real-time voice AI latency budget

Voir ce sujet

One-second turn deadline

A production voice agent must begin speaking within about one second of the user's utterance ending, or the exchange feels broken.

Éléments favorables

Real-Time Voice AI Stack for Agents: Architecture Guide

Extrait original

It has about a second to start talking before the exchange feels broken.
Contexte

A production voice agent chains transport, speech recognition, a language model, speech synthesis, and orchestration into a single real-time exchange, and each layer draws from the same shared clock.

Jose Nicholas Francisco
Partager un aperçu

perceived latency thresholds

Voir ce sujet

600–800 ms perceived willingness drop

Perceived willingness to continue the interaction begins to decline after 600 ms and drops significantly between 700 and 800 ms.

Éléments favorables

Real-Time Voice AI Stack for Agents: Architecture Guide

Extrait original

Perceived willingness begins to drop after 600 ms and steps down significantly from 700 to 800 ms.
Contexte

Transport, recognition, endpointing, the LLM's first token, TTS first byte, and the return network hop each take a slice of that budget. This guide breaks down what each layer costs and where you can cut to keep your voice AI stack inside it.

Jose Nicholas Francisco
Partager un aperçu

production latency SLOs

Voir ce sujet

P95 under 1,000 ms target

A realistic production per-turn latency target is P95 under approximately 1,000 ms; beyond a couple of seconds, callers begin to disengage, and percentile-based SLOs—not averages—should be used and calibrated against user studies and real caller feedback.

Éléments favorables

Real-Time Voice AI Stack for Agents: Architecture Guide

Extrait original

Target a P95 turn latency under about 1,000 ms, and treat anything beyond a couple of seconds as the point where callers start to disengage. Set percentile SLOs, not averages, and calibrate against user-study research like this IEEE study on perceived response delay in voice interfaces, then confirm the specific thresholds against your own caller feedback.
Jose Nicholas Francisco
Partager un aperçu

voice agent function calling

Voir ce sujet

Function calling enables direct answers instead of call transfers

Function calling allows a voice agent to directly answer questions like 'Where’s my order?' rather than transferring the caller to a human agent.

Éléments favorables

Voice Agent Function Calling: Real-Time Data Lookup Guide

Extrait original

Function calling is how a voice agent answers "where's my order?" instead of transferring the caller to someone who can.
Contexte

One order-status lookup on the inbound telephony agent pattern takes two round trips, three function calls, and a spoken line for every way it can fail.

Jose Nicholas Francisco
Partager un aperçu

latency tolerance in voice interactions

Voir ce sujet

Perceived willingness drops significantly after 600–800 ms of silence

Ratings of perceived willingness drop noticeably at 600 ms and become statistically significant between 700 and 800 ms, based on the 2013 gap-tolerance study by Roberts and Francis.

Éléments favorables

Voice Agent Function Calling: Real-Time Data Lookup Guide

Extrait original

Ratings of perceived willingness drop off noticeably at 600ms in the classic 2013 gap-tolerance study from Roberts and Francis, and the difference turns statistically significant between 700 and 800ms.
Contexte

Dead air is what the caller hears while your handler works, and it costs you fast. InjectAgentMessage is what covers that silence in the reference pattern.

Jose Nicholas Francisco
Partager un aperçu

function selection accuracy

Voir ce sujet

Tool count degrades selection accuracy: 98% at 10 tools → 88% at 100

The HumanMCP study measured a decline in function selection accuracy from 98% with 10 tools to 88% with 100 tools.

Éléments favorables

Voice Agent Function Calling: Real-Time Data Lookup Guide

Extrait original

The HumanMCP study measured a drop from 98% at 10 tools to 88% at 100.
Contexte

When the toolset grows, selection accuracy falls. Add functions one at a time and re-check selection accuracy after each. One lookup, three failure lines, and a filler function make a complete first release.

Jose Nicholas Francisco
Partager un aperçu

session limits in voice agents

Voir ce sujet

Two-hour session limit applies regardless of in-flight lookups

Voice agent sessions expire after two hours; an in-flight function call does not extend the session, and the system emits MAXIMUM_SESSION_LENGTH_APPROACHING at 1h 55m and MAXIMUM_SESSION_LENGTH_REACHED at 2h.

Éléments favorables

Voice Agent Function Calling: Real-Time Data Lookup Guide

Extrait original

Yes, at two hours, and an in-flight lookup won't hold the session open. You get a Warning with MAXIMUM_SESSION_LENGTH_APPROACHING at 1 hour 55 minutes, then an Error with MAXIMUM_SESSION_LENGTH_REACHED at the two-hour limit .
Contexte

KeepAlive doesn't extend it, and the docs don't cover a pending FunctionCallRequest at close. Finish or abandon it, then reconnect with agent.context .

Jose Nicholas Francisco
Partager un aperçu

Propos par date de source2

Les propos sont classés par date de publication de la source originale ; une différence de formulation ne prouve pas un changement de position.