TTFA 比较采用受控流程
作者描述了一项 TTFA 对比,使用批大小 1、相同的 50 条英语 CV3-Eval 提示、相同硬件及各模型默认声音。前 3 次运行作为预热被舍弃,报告中位数。流式 TTFA 测量首个音频片段;非流式 TTFA 等待整句语音。
支持这项说法
Open TTS Leaderboard:多语言文本转语音与声音克隆的可扩展评估
对于流式模型(“Streaming API”下标有 ✅),TTFA 指首个音频片段到达所需的时间。对于非流式模型,TTFA 指整句生成完毕所需的时间,因为在此之前无法开始播放。每个模型每次运行一段音频(批大小 1),在同一硬件上使用来自 CV3-Eval 的相同 50 条英语提示,并以其默认声音运行。我们舍弃前 3 次运行作为预热,并报告其余运行的 TTFA 中位值。
原始摘录
For streaming models (✅ under “Streaming API”) it's the time until the first audio chunk arrives. For non-streaming models, it's the time until the whole utterance is generated, because playback can't start any earlier. Every model runs one audio at a time (batch size 1), on the same 50 English prompts from CV3-Eval, on the same hardware and in its default voice. We drop the first 3 runs as warm-up and report the median TTFA across the rest.