Open TTS Leaderboard:多语言文本转语音与声音克隆的可扩展评估

Hugging Face Blog ·

Open TTS Leaderboard 使用客观指标比较文本转语音和声音克隆模型。作者表示,这些指标补充人类偏好排名。他们提醒,英语表现不一定能迁移到其他语言。其首音频延迟(time-to-first-audio)对比采用批大小为 1,在同一硬件上使用相同的 50 条英语提示,并采用各模型的默认声音。舍弃前 3 次运行,报告中位数。 阅读 3 条观点,查看支持证据与原始来源。

Eric Bezzam, Steven Zheng, Eustache Le Bihan, mrfakename

理解这篇

3 个要点

综合解读

  1. 客观指标补充人类偏好排名

    作者指出,客观指标是对人类偏好排名的补充。基于ASR的WER是可懂度的代理指标,说话人相似度用于估计声音身份保留程度;二者均不直接衡量自然度、表现力或听者偏好。

    支持这项说法 1

    重要的是,Open TTS Leaderboard 并不能取代人类偏好排名。基于 ASR 的 WER 仅为可懂度的代理指标,说话人相似度则用于估计声音身份的保留程度。两者均不直接衡量自然度、表现力或听者偏好。尽管如此,它们仍可为基于投票的排行榜提供参考,以确定应将哪些模型纳入其评估。

    Eric Bezzam, Steven Zheng, Eustache Le Bihan, mrfakename · 段落 11

    原始摘录
    Importantly, the Open TTS Leaderboard does not replace human preference ranking. ASR-based WER provides a proxy for intelligibility, while speaker similarity estimates voice identity preservation. Neither directly measures naturalness, expressiveness, or listener preference. Nevertheless, they can even inform voting-based leaderboards which models to include in their evaluations.
    回到原文语境 →

    继续探索

    TTS 评估 →
  2. 英语得分不能代表多语言表现

    作者指出,英语表现不一定能迁移到其他语言。该排行榜报告中文、日文和韩文的字符错误率(CER),并计算这三种语言的宏平均值。Seed TTS Eval 用于英语和中文;其余语言使用 CV3 Eval(零样本)。

    支持这项说法 1

    英语上的表现不一定能迁移到其他语言。支持切换多种语言,以便按多语言表现对模型进行排名。Seed TTS Eval 仅提供英语和中文音频,因此其他语言直接采用 CV3 Eval(零样本)的得分。注意:中文、日文和韩文是基于字符的语言,报告的是字符错误率(CER);跨语言的‘平均 WER’为各语言 WER 的宏平均。

    Eric Bezzam, Steven Zheng, Eustache Le Bihan, mrfakename · 段落 16

    原始摘录
    English performance doesn't necessarily translate to other languages. Multiple languages can be toggled to rank models on multilingual performance. Seed TTS Eval only has audio for English and Chinese, so the other languages are simply the score on CV3 Eval (zero shot). Note that Chinese, Japanese, and Korean are character-based languages and so character error rate (CER) is reported, and the “Average WER” across languages is a macro-average across languages.
    回到原文语境 →

    继续探索

    多语言 TTS →
  3. TTFA 比较采用受控流程

    作者描述了一项 TTFA 对比,使用批大小 1、相同的 50 条英语 CV3-Eval 提示、相同硬件及各模型默认声音。前 3 次运行作为预热被舍弃,报告中位数。流式 TTFA 测量首个音频片段;非流式 TTFA 等待整句语音。

    支持这项说法 1

    对于流式模型(“Streaming API”下标有 ✅),TTFA 指首个音频片段到达所需的时间。对于非流式模型,TTFA 指整句生成完毕所需的时间,因为在此之前无法开始播放。每个模型每次运行一段音频(批大小 1),在同一硬件上使用来自 CV3-Eval 的相同 50 条英语提示,并以其默认声音运行。我们舍弃前 3 次运行作为预热,并报告其余运行的 TTFA 中位值。

    Eric Bezzam, Steven Zheng, Eustache Le Bihan, mrfakename · 段落 28

    原始摘录
    For streaming models (✅ under “Streaming API”) it's the time until the first audio chunk arrives. For non-streaming models, it's the time until the whole utterance is generated, because playback can't start any earlier. Every model runs one audio at a time (batch size 1), on the same 50 English prompts from CV3-Eval, on the same hardware and in its default voice. We drop the first 3 runs as warm-up and report the median TTFA across the rest.
    回到原文语境 →

    继续探索

    语音延迟 →

关键段落3

带明确归属与语境的原文片段。打开原始文本核查出处。

多语言 TTS

英语得分不能代表多语言表现

英语上的表现不一定能迁移到其他语言。支持切换多种语言,以便按多语言表现对模型进行排名。Seed TTS Eval 仅提供英语和中文音频,因此其他语言直接采用 CV3 Eval(零样本)的得分。注意:中文、日文和韩文是基于字符的语言,报告的是字符错误率(CER);跨语言的‘平均 WER’为各语言 WER 的宏平均。

原始摘录
English performance doesn't necessarily translate to other languages. Multiple languages can be toggled to rank models on multilingual performance. Seed TTS Eval only has audio for English and Chinese, so the other languages are simply the score on CV3 Eval (zero shot). Note that Chinese, Japanese, and Korean are character-based languages and so character error rate (CER) is reported, and the “Average WER” across languages is a macro-average across languages.
TTS 评估

客观指标补充人类偏好排名

重要的是,Open TTS Leaderboard 并不能取代人类偏好排名。基于 ASR 的 WER 仅为可懂度的代理指标,说话人相似度则用于估计声音身份的保留程度。两者均不直接衡量自然度、表现力或听者偏好。尽管如此,它们仍可为基于投票的排行榜提供参考,以确定应将哪些模型纳入其评估。

原始摘录
Importantly, the Open TTS Leaderboard does not replace human preference ranking. ASR-based WER provides a proxy for intelligibility, while speaker similarity estimates voice identity preservation. Neither directly measures naturalness, expressiveness, or listener preference. Nevertheless, they can even inform voting-based leaderboards which models to include in their evaluations.
语音延迟

TTFA 比较采用受控流程

对于流式模型(“Streaming API”下标有 ✅),TTFA 指首个音频片段到达所需的时间。对于非流式模型,TTFA 指整句生成完毕所需的时间,因为在此之前无法开始播放。每个模型每次运行一段音频(批大小 1),在同一硬件上使用来自 CV3-Eval 的相同 50 条英语提示,并以其默认声音运行。我们舍弃前 3 次运行作为预热,并报告其余运行的 TTFA 中位值。

原始摘录
For streaming models (✅ under “Streaming API”) it's the time until the first audio chunk arrives. For non-streaming models, it's the time until the whole utterance is generated, because playback can't start any earlier. Every model runs one audio at a time (batch size 1), on the same 50 English prompts from CV3-Eval, on the same hardware and in its default voice. We drop the first 3 runs as warm-up and report the median TTFA across the rest.

这里提到的

全部提及对象

Open TTS Leaderboard

仅提及

作者澄清,Open TTS排行榜使用ASR词错误率(WER)和说话人相似度等客观指标作为可懂度和语音身份保持性的代理指标,而非直接测量,并不取代人类偏好排序。

查看支持证据 · Eric Bezzam, Steven Zheng, Eustache Le Bihan, mrfakename

来源与研究方法

这些观点均关联原始来源。转述已明确标注,不作为逐字原话展示。

打开转录或来源材料 (在新标签页中打开)报告问题