Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning

Hugging Face Blog ·

The Open TTS Leaderboard compares text-to-speech and voice-cloning models using objective metrics. The authors say these metrics complement human preference rankings. They caution that English performance does not necessarily carry over to other languages. Its time-to-first-audio comparison uses batch size 1 and the same 50 English prompts on the same hardware, with each model’s default voice. The first 3 runs are discarded and the median is reported. Read 3 viewpoints with supporting evidence and source links.

Eric Bezzam, Steven Zheng, Eustache Le Bihan, mrfakename

Understand this piece

3 key points

Synthesis

  1. Objective metrics complement human preference

    The authors state that objective metrics complement human preference rankings. ASR-based WER is a proxy for intelligibility and speaker similarity estimates voice identity preservation; neither directly measures naturalness, expressiveness, or listener preference.

    Supporting evidence 1

    Original excerpt

    Importantly, the Open TTS Leaderboard does not replace human preference ranking. ASR-based WER provides a proxy for intelligibility, while speaker similarity estimates voice identity preservation. Neither directly measures naturalness, expressiveness, or listener preference. Nevertheless, they can even inform voting-based leaderboards which models to include in their evaluations.

    Eric Bezzam, Steven Zheng, Eustache Le Bihan, mrfakename · Paragraph 11

    Read in source context →

    Continue exploring

    TTS Evaluation →
  2. English scores do not establish multilingual performance

    The authors caution that English performance does not necessarily carry over to other languages. The leaderboard reports CER for Chinese, Japanese, and Korean, and computes a macro-average across languages. Seed TTS Eval covers English and Chinese; other languages use CV3 Eval (zero shot).

    Supporting evidence 1

    Original excerpt

    English performance doesn't necessarily translate to other languages. Multiple languages can be toggled to rank models on multilingual performance. Seed TTS Eval only has audio for English and Chinese, so the other languages are simply the score on CV3 Eval (zero shot). Note that Chinese, Japanese, and Korean are character-based languages and so character error rate (CER) is reported, and the “Average WER” across languages is a macro-average across languages.

    Eric Bezzam, Steven Zheng, Eustache Le Bihan, mrfakename · Paragraph 16

    Read in source context →

    Continue exploring

    Multilingual TTS →
  3. TTFA comparisons use a controlled protocol

    The authors describe a TTFA comparison using batch size 1, the same 50 English CV3-Eval prompts, the same hardware, and each model’s default voice. The first 3 runs are discarded as warm-up and the median is reported. Streaming TTFA measures the first audio chunk; non-streaming TTFA waits for the whole utterance.

    Supporting evidence 1

    Original excerpt

    For streaming models (✅ under “Streaming API”) it's the time until the first audio chunk arrives. For non-streaming models, it's the time until the whole utterance is generated, because playback can't start any earlier. Every model runs one audio at a time (batch size 1), on the same 50 English prompts from CV3-Eval, on the same hardware and in its default voice. We drop the first 3 runs as warm-up and report the median TTFA across the rest.

    Eric Bezzam, Steven Zheng, Eustache Le Bihan, mrfakename · Paragraph 28

    Read in source context →

    Continue exploring

    Speech Latency →

Key passages3

Attributed passages with the context to verify them. Open the original text to check the source.

Multilingual TTS

English scores do not establish multilingual performance

Original excerpt

English performance doesn't necessarily translate to other languages. Multiple languages can be toggled to rank models on multilingual performance. Seed TTS Eval only has audio for English and Chinese, so the other languages are simply the score on CV3 Eval (zero shot). Note that Chinese, Japanese, and Korean are character-based languages and so character error rate (CER) is reported, and the “Average WER” across languages is a macro-average across languages.
TTS Evaluation

Objective metrics complement human preference

Original excerpt

Importantly, the Open TTS Leaderboard does not replace human preference ranking. ASR-based WER provides a proxy for intelligibility, while speaker similarity estimates voice identity preservation. Neither directly measures naturalness, expressiveness, or listener preference. Nevertheless, they can even inform voting-based leaderboards which models to include in their evaluations.
Speech Latency

TTFA comparisons use a controlled protocol

Original excerpt

For streaming models (✅ under “Streaming API”) it's the time until the first audio chunk arrives. For non-streaming models, it's the time until the whole utterance is generated, because playback can't start any earlier. Every model runs one audio at a time (batch size 1), on the same 50 English prompts from CV3-Eval, on the same hardware and in its default voice. We drop the first 3 runs as warm-up and report the median TTFA across the rest.

Mentioned here

All mentioned things

Open TTS Leaderboard

Mention only

The authors clarify that the Open TTS Leaderboard uses objective metrics like ASR-based WER and speaker similarity as proxies—not direct measures—of intelligibility and voice identity preservation, and does not replace human preference ranking.

Read supporting evidence · Eric Bezzam, Steven Zheng, Eustache Le Bihan, mrfakename

Source & methodology

These viewpoints are linked to their original sources. Paraphrases are labeled and are not verbatim quotes.

Open transcript or source material (opens in a new tab)Report an issue