Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning

Hugging Face Blog ·

The Open TTS Leaderboard compares text-to-speech and voice-cloning models using objective metrics. The authors say these metrics complement human preference rankings. They caution that English performance does not necessarily carry over to other languages. Its time-to-first-audio comparison uses batch size 1 and the same 50 English prompts on the same hardware, with each model’s default voice. The first 3 runs are discarded and the median is reported. Lisez 3 points de vue avec leurs éléments à l’appui et les liens vers les sources.

Eric Bezzam, Steven Zheng, Eustache Le Bihan, mrfakename

En un coup d’œil

  • Objective metrics complement human preference

    The authors state that objective metrics complement human preference rankings. ASR-based WER is a proxy for intelligibility and speaker similarity estimates voice identity preservation; neither directly measures naturalness, expressiveness, or listener preference.

    Lire le moment probant · Paragraphe 11
  • English scores do not establish multilingual performance

    The authors caution that English performance does not necessarily carry over to other languages. The leaderboard reports CER for Chinese, Japanese, and Korean, and computes a macro-average across languages. Seed TTS Eval covers English and Chinese; other languages use CV3 Eval (zero shot).

    Lire le moment probant · Paragraphe 16
  • TTFA comparisons use a controlled protocol

    The authors describe a TTFA comparison using batch size 1, the same 50 English CV3-Eval prompts, the same hardware, and each model’s default voice. The first 3 runs are discarded as warm-up and the median is reported. Streaming TTFA measures the first audio chunk; non-streaming TTFA waits for the whole utterance.

    Lire le moment probant · Paragraphe 28

Passages clés3

Passages attribués et accompagnés du contexte nécessaire à leur vérification. Ouvrez le texte original pour vérifier la source.

Multilingual TTS

English scores do not establish multilingual performance

Extrait original

English performance doesn't necessarily translate to other languages. Multiple languages can be toggled to rank models on multilingual performance. Seed TTS Eval only has audio for English and Chinese, so the other languages are simply the score on CV3 Eval (zero shot). Note that Chinese, Japanese, and Korean are character-based languages and so character error rate (CER) is reported, and the “Average WER” across languages is a macro-average across languages.
TTS Evaluation

Objective metrics complement human preference

Extrait original

Importantly, the Open TTS Leaderboard does not replace human preference ranking. ASR-based WER provides a proxy for intelligibility, while speaker similarity estimates voice identity preservation. Neither directly measures naturalness, expressiveness, or listener preference. Nevertheless, they can even inform voting-based leaderboards which models to include in their evaluations.
Speech Latency

TTFA comparisons use a controlled protocol

Extrait original

For streaming models (✅ under “Streaming API”) it's the time until the first audio chunk arrives. For non-streaming models, it's the time until the whole utterance is generated, because playback can't start any earlier. Every model runs one audio at a time (batch size 1), on the same 50 English prompts from CV3-Eval, on the same hardware and in its default voice. We drop the first 3 runs as warm-up and report the median TTFA across the rest.

Source et méthodologie

Ces points de vue renvoient à leurs sources originales. Les reformulations sont signalées et ne sont pas des citations mot à mot.

Ouvrir la transcription ou les documents sources (s’ouvre dans un nouvel onglet)Signaler un problème