Voice Agent Function Calling: Real-Time Data Lookup Guide

Deepgram Learn ·

A guide on voice agent function calling that covers direct answering capability, latency tolerance thresholds from a 2013 study, function selection accuracy decline with increasing tool count, and fixed two-hour session limits including warning and error signals. Read 4 viewpoints with supporting evidence and source links.

Understand this piece

4 key points

Synthesis

  1. Function calling enables direct answers instead of call transfers

    Function calling allows a voice agent to directly answer questions like 'Where’s my order?' rather than transferring the caller to a human agent.

    Supporting evidence 1

    Original excerpt

    Function calling is how a voice agent answers "where's my order?" instead of transferring the caller to someone who can.

    Jose Nicholas Francisco · Paragraph 1

    Context

    One order-status lookup on the inbound telephony agent pattern takes two round trips, three function calls, and a spoken line for every way it can fail.

    Read in source context →
  2. Perceived willingness drops significantly after 600–800 ms of silence

    Ratings of perceived willingness drop noticeably at 600 ms and become statistically significant between 700 and 800 ms, based on the 2013 gap-tolerance study by Roberts and Francis.

    Supporting evidence 1

    Original excerpt

    Ratings of perceived willingness drop off noticeably at 600ms in the classic 2013 gap-tolerance study from Roberts and Francis, and the difference turns statistically significant between 700 and 800ms.

    Jose Nicholas Francisco · Paragraph 22

    Context

    Dead air is what the caller hears while your handler works, and it costs you fast. InjectAgentMessage is what covers that silence in the reference pattern.

    Read in source context →
  3. Tool count degrades selection accuracy: 98% at 10 tools → 88% at 100

    The HumanMCP study measured a decline in function selection accuracy from 98% with 10 tools to 88% with 100 tools.

    Supporting evidence 1

    Original excerpt

    The HumanMCP study measured a drop from 98% at 10 tools to 88% at 100.

    Jose Nicholas Francisco · Paragraph 72

    Context

    When the toolset grows, selection accuracy falls. Add functions one at a time and re-check selection accuracy after each. One lookup, three failure lines, and a filler function make a complete first release.

    Read in source context →
  4. Two-hour session limit applies regardless of in-flight lookups

    Voice agent sessions expire after two hours; an in-flight function call does not extend the session, and the system emits MAXIMUM_SESSION_LENGTH_APPROACHING at 1h 55m and MAXIMUM_SESSION_LENGTH_REACHED at 2h.

    Supporting evidence 1

    Original excerpt

    Yes, at two hours, and an in-flight lookup won't hold the session open. You get a Warning with MAXIMUM_SESSION_LENGTH_APPROACHING at 1 hour 55 minutes, then an Error with MAXIMUM_SESSION_LENGTH_REACHED at the two-hour limit .

    Jose Nicholas Francisco · Paragraph 83

    Context

    KeepAlive doesn't extend it, and the docs don't cover a pending FunctionCallRequest at close. Finish or abandon it, then reconnect with agent.context .

    Read in source context →

Key passages4

Attributed passages with the context to verify them. Open the original text to check the source.

function selection accuracy

Tool count degrades selection accuracy: 98% at 10 tools → 88% at 100

Original excerpt

The HumanMCP study measured a drop from 98% at 10 tools to 88% at 100.
Context

When the toolset grows, selection accuracy falls. Add functions one at a time and re-check selection accuracy after each. One lookup, three failure lines, and a filler function make a complete first release.

voice agent function calling

Function calling enables direct answers instead of call transfers

Original excerpt

Function calling is how a voice agent answers "where's my order?" instead of transferring the caller to someone who can.
Context

One order-status lookup on the inbound telephony agent pattern takes two round trips, three function calls, and a spoken line for every way it can fail.

latency tolerance in voice interactions

Perceived willingness drops significantly after 600–800 ms of silence

Original excerpt

Ratings of perceived willingness drop off noticeably at 600ms in the classic 2013 gap-tolerance study from Roberts and Francis, and the difference turns statistically significant between 700 and 800ms.
Context

Dead air is what the caller hears while your handler works, and it costs you fast. InjectAgentMessage is what covers that silence in the reference pattern.

session limits in voice agents

Two-hour session limit applies regardless of in-flight lookups

Original excerpt

Yes, at two hours, and an in-flight lookup won't hold the session open. You get a Warning with MAXIMUM_SESSION_LENGTH_APPROACHING at 1 hour 55 minutes, then an Error with MAXIMUM_SESSION_LENGTH_REACHED at the two-hour limit .
Context

KeepAlive doesn't extend it, and the docs don't cover a pending FunctionCallRequest at close. Finish or abandon it, then reconnect with agent.context .

Source & methodology

These viewpoints are linked to their original sources. Paraphrases are labeled and are not verbatim quotes.

Open transcript or source material (opens in a new tab)Report an issue

Explore these viewpoints by person