Cartesia | A Guide to Choosing Voice AI Models

Cartesia Blog ·

Zubin Pratap explains how to choose voice AI models for their intended use cases, assess latency and turn detection, and evaluate word error rate across relevant error categories. Read 4 viewpoints with supporting evidence and source links.

Understand this piece

4 key points

Synthesis

  1. Select models by use-case performance, not lab benchmarks

    Voice AI models must be selected, analyzed, and measured against their performance for their intended use case—not based on laboratory-condition benchmarks, because real-life conversations are messy and enterprise-grade voice agents are a hard engineering problem.

    Supporting evidence 1

    Original excerpt

    It is tempting to choose the models based on benchmarks that measure them in laboratory conditions. But real-life conversations are messy, and that makes enterprise-grade voice agents a hard engineering problem. We cannot solve these hard problems by treating voice models as interchangeable commodities. Instead, they must be picked, analysed and measured against their performance for their intended use case .

    Zubin Pratap · Paragraph 4

    Read in source context →

    Continue exploring

    model evaluation →
  2. TTFS and TTCT as voice latency metrics

    The author considers TTFS (Time to Final Segment), also called TTCT (Time to Complete Transcript), the metric that matters for voice agents.

    Supporting evidence 1

    Original excerpt

    For voice agents, we believe the metric that matters is TTFS - Time to Final Segment (aka the Time to Complete Transcript (TTCT)).

    Zubin Pratap · Paragraph 14

    Read in source context →

    Continue exploring

    model evaluation →
  3. Turn detection is critical and distinct from ASR accuracy

    Turn detection, also called end pointing, detects when a user has finished their turn. The author describes it as extremely hard without visual or other human conversational cues, and says poor turn detection causes long pauses and high latency.

    Supporting evidence 1

    Original excerpt

    Turn-detection is also known as end pointing – because it refers to detecting when the user has finished their turn. It is extremely hard to do accurately as Speech to Text models do not have visual or other cues like humans do in conversation. That makes high-quality turn detection incredibly important for conversational AI - poor turn detection results in very long awkward pauses and high latency in the voice agent’s conversation.

    Zubin Pratap · Paragraph 18

    Read in source context →

    Continue exploring

    model evaluation →
  4. WER is misleading without category-level analysis

    Word Error Rate can mislead when the error classes are irrelevant or when a non-streaming ASR model is compared with a streaming model. The author says many benchmarks fail to distinguish streaming and non-streaming use cases.

    Supporting evidence 1

    Original excerpt

    Word Error Rate (WER) for STT/ASR is a misleading metric if you’re looking at the wrong or irrelevant classes of errors. Or if you’re comparing a non-streaming ASR model (which is inherently easier to be accurate on) with a streaming ASR model. Unfortunately many benchmarks do not distinguish between ASR streaming vs non-streaming use cases.

    Zubin Pratap · Paragraph 25

    Read in source context →

    Continue exploring

    model evaluation →

Key passages4

Attributed passages with the context to verify them. Open the original text to check the source.

model evaluation

TTFS and TTCT as voice latency metrics

Original excerpt

For voice agents, we believe the metric that matters is TTFS - Time to Final Segment (aka the Time to Complete Transcript (TTCT)).
model evaluation

Turn detection is critical and distinct from ASR accuracy

Original excerpt

Turn-detection is also known as end pointing – because it refers to detecting when the user has finished their turn. It is extremely hard to do accurately as Speech to Text models do not have visual or other cues like humans do in conversation. That makes high-quality turn detection incredibly important for conversational AI - poor turn detection results in very long awkward pauses and high latency in the voice agent’s conversation.
model evaluation

Select models by use-case performance, not lab benchmarks

Original excerpt

It is tempting to choose the models based on benchmarks that measure them in laboratory conditions. But real-life conversations are messy, and that makes enterprise-grade voice agents a hard engineering problem. We cannot solve these hard problems by treating voice models as interchangeable commodities. Instead, they must be picked, analysed and measured against their performance for their intended use case .
model evaluation

WER is misleading without category-level analysis

Original excerpt

Word Error Rate (WER) for STT/ASR is a misleading metric if you’re looking at the wrong or irrelevant classes of errors. Or if you’re comparing a non-streaming ASR model (which is inherently easier to be accurate on) with a streaming ASR model. Unfortunately many benchmarks do not distinguish between ASR streaming vs non-streaming use cases.

Source & methodology

These viewpoints are linked to their original sources. Paraphrases are labeled and are not verbatim quotes.

Open transcript or source material (opens in a new tab)Report an issue

Explore these viewpoints by person

Continue with this topic

More sources on topics discussed here. Shared topics do not imply agreement.