TTFA comparisons use a controlled protocol
The authors describe a TTFA comparison using batch size 1, the same 50 English CV3-Eval prompts, the same hardware, and each model’s default voice. The first 3 runs are discarded as warm-up and the median is reported. Streaming TTFA measures the first audio chunk; non-streaming TTFA waits for the whole utterance.
Evidencia a favor
Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
Extracto original
For streaming models (✅ under “Streaming API”) it's the time until the first audio chunk arrives. For non-streaming models, it's the time until the whole utterance is generated, because playback can't start any earlier. Every model runs one audio at a time (batch size 1), on the same 50 English prompts from CV3-Eval, on the same hardware and in its default voice. We drop the first 3 runs as warm-up and report the median TTFA across the rest.