The authors clarify that the Open TTS Leaderboard uses objective metrics like ASR-based WER and speaker similarity as proxies—not direct measures—of intelligibility and voice identity preservation, and does not replace human preference ranking.
Eric Bezzam, Steven Zheng, Eustache Le Bihan, mrfakename ·
Supporting evidence
Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
Original excerpt
Importantly, the Open TTS Leaderboard does not replace human preference ranking. ASR-based WER provides a proxy for intelligibility, while speaker similarity estimates voice identity preservation. Neither directly measures naturalness, expressiveness, or listener preference. Nevertheless, they can even inform voting-based leaderboards which models to include in their evaluations.
About this interpretation
The object-specific interpretation and Chinese translation were checked independently against the source. This is an AI semantic review, not playback verification. Reviewed Oct 5, 2026 · qwen3.8-max-0902
Report an issue