Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning

Hugging Face Blog ·

The Open TTS Leaderboard compares text-to-speech and voice-cloning models using objective metrics. The authors say these metrics complement human preference rankings. They caution that English performance does not necessarily carry over to other languages. Its time-to-first-audio comparison uses batch size 1 and the same 50 English prompts on the same hardware, with each model’s default voice. The first 3 runs are discarded and the median is reported. Lies 3 Standpunkte mit Belegen und Links zu den Originalquellen.

Eric Bezzam, Steven Zheng, Eustache Le Bihan, mrfakename

Auf einen Blick

  • Objective metrics complement human preference

    The authors state that objective metrics complement human preference rankings. ASR-based WER is a proxy for intelligibility and speaker similarity estimates voice identity preservation; neither directly measures naturalness, expressiveness, or listener preference.

    Unterstützendes Moment lesen · Absatz 11
  • English scores do not establish multilingual performance

    The authors caution that English performance does not necessarily carry over to other languages. The leaderboard reports CER for Chinese, Japanese, and Korean, and computes a macro-average across languages. Seed TTS Eval covers English and Chinese; other languages use CV3 Eval (zero shot).

    Unterstützendes Moment lesen · Absatz 16
  • TTFA comparisons use a controlled protocol

    The authors describe a TTFA comparison using batch size 1, the same 50 English CV3-Eval prompts, the same hardware, and each model’s default voice. The first 3 runs are discarded as warm-up and the median is reported. Streaming TTFA measures the first audio chunk; non-streaming TTFA waits for the whole utterance.

    Unterstützendes Moment lesen · Absatz 28

Wichtige Passagen3

Zugeordnete Passagen mit dem Kontext zur Überprüfung. Öffnen Sie den Originaltext, um die Quelle zu prüfen.

Multilingual TTS

English scores do not establish multilingual performance

Originalauszug

English performance doesn't necessarily translate to other languages. Multiple languages can be toggled to rank models on multilingual performance. Seed TTS Eval only has audio for English and Chinese, so the other languages are simply the score on CV3 Eval (zero shot). Note that Chinese, Japanese, and Korean are character-based languages and so character error rate (CER) is reported, and the “Average WER” across languages is a macro-average across languages.
TTS Evaluation

Objective metrics complement human preference

Originalauszug

Importantly, the Open TTS Leaderboard does not replace human preference ranking. ASR-based WER provides a proxy for intelligibility, while speaker similarity estimates voice identity preservation. Neither directly measures naturalness, expressiveness, or listener preference. Nevertheless, they can even inform voting-based leaderboards which models to include in their evaluations.
Speech Latency

TTFA comparisons use a controlled protocol

Originalauszug

For streaming models (✅ under “Streaming API”) it's the time until the first audio chunk arrives. For non-streaming models, it's the time until the whole utterance is generated, because playback can't start any earlier. Every model runs one audio at a time (batch size 1), on the same 50 English prompts from CV3-Eval, on the same hardware and in its default voice. We drop the first 3 runs as warm-up and report the median TTFA across the rest.

Quelle & Methodik

Diese Standpunkte sind mit ihren Originalquellen verknüpft. Paraphrasen sind gekennzeichnet und keine wörtlichen Zitate.

Transkript oder Quellenmaterial öffnen (wird in einem neuen Tab geöffnet)Ein Problem melden