A TOPIC, IN CONTEXT

TTS Evaluation

Judgments in this source concerning TTS Evaluation. Explore 1 viewpoint with evidence from 1 source.

0 people · 1 sources · 1 viewpoints

Content updated:

Explore connections ↗

Viewpoint map

0 people · 1 sources · 1 viewpoints

Objective metrics complement human preference

The authors state that objective metrics complement human preference rankings. ASR-based WER is a proxy for intelligibility and speaker similarity estimates voice identity preservation; neither directly measures naturalness, expressiveness, or listener preference.

Supporting evidence

Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning

Original excerpt

Importantly, the Open TTS Leaderboard does not replace human preference ranking. ASR-based WER provides a proxy for intelligibility, while speaker similarity estimates voice identity preservation. Neither directly measures naturalness, expressiveness, or listener preference. Nevertheless, they can even inform voting-based leaderboards which models to include in their evaluations.

These findings reflect the available sources, not an exhaustive or current view.