Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning

Hugging Face Blog ·

The Open TTS Leaderboard compares text-to-speech and voice-cloning models using objective metrics. The authors say these metrics complement human preference rankings. They caution that English performance does not necessarily carry over to other languages. Its time-to-first-audio comparison uses batch size 1 and the same 50 English prompts on the same hardware, with each model’s default voice. The first 3 runs are discarded and the median is reported. Lee 3 puntos de vista con sus evidencias y enlaces a las fuentes.

Eric Bezzam, Steven Zheng, Eustache Le Bihan, mrfakename

De un vistazo

  • Objective metrics complement human preference

    The authors state that objective metrics complement human preference rankings. ASR-based WER is a proxy for intelligibility and speaker similarity estimates voice identity preservation; neither directly measures naturalness, expressiveness, or listener preference.

    Ver el momento de apoyo · Párrafo 11
  • English scores do not establish multilingual performance

    The authors caution that English performance does not necessarily carry over to other languages. The leaderboard reports CER for Chinese, Japanese, and Korean, and computes a macro-average across languages. Seed TTS Eval covers English and Chinese; other languages use CV3 Eval (zero shot).

    Ver el momento de apoyo · Párrafo 16
  • TTFA comparisons use a controlled protocol

    The authors describe a TTFA comparison using batch size 1, the same 50 English CV3-Eval prompts, the same hardware, and each model’s default voice. The first 3 runs are discarded as warm-up and the median is reported. Streaming TTFA measures the first audio chunk; non-streaming TTFA waits for the whole utterance.

    Ver el momento de apoyo · Párrafo 28

Pasajes clave3

Pasajes atribuidos con contexto para verificarlos. Abra el texto original para comprobar la fuente.

Multilingual TTS

English scores do not establish multilingual performance

Extracto original

English performance doesn't necessarily translate to other languages. Multiple languages can be toggled to rank models on multilingual performance. Seed TTS Eval only has audio for English and Chinese, so the other languages are simply the score on CV3 Eval (zero shot). Note that Chinese, Japanese, and Korean are character-based languages and so character error rate (CER) is reported, and the “Average WER” across languages is a macro-average across languages.
TTS Evaluation

Objective metrics complement human preference

Extracto original

Importantly, the Open TTS Leaderboard does not replace human preference ranking. ASR-based WER provides a proxy for intelligibility, while speaker similarity estimates voice identity preservation. Neither directly measures naturalness, expressiveness, or listener preference. Nevertheless, they can even inform voting-based leaderboards which models to include in their evaluations.
Speech Latency

TTFA comparisons use a controlled protocol

Extracto original

For streaming models (✅ under “Streaming API”) it's the time until the first audio chunk arrives. For non-streaming models, it's the time until the whole utterance is generated, because playback can't start any earlier. Every model runs one audio at a time (batch size 1), on the same 50 English prompts from CV3-Eval, on the same hardware and in its default voice. We drop the first 3 runs as warm-up and report the median TTFA across the rest.

Fuente y metodología

Estas perspectivas enlazan a sus fuentes originales. Las paráfrasis están identificadas y no son citas textuales.

Abrir transcripción o material de origen (se abre en una pestaña nueva)Reportar un problema