Temas / model evaluationPunto de vista parafraseado
WER is misleading without category-level analysis
Word Error Rate can mislead when the error classes are irrelevant or when a non-streaming ASR model is compared with a streaming model. The author says many benchmarks fail to distinguish streaming and non-streaming use cases.
Detrás de la opinión
Las traducciones son para facilitar la lectura; los extractos originales siguen siendo la source evidence.
Cartesia | A Guide to Choosing Voice AI Models
Extracto original
Word Error Rate (WER) for STT/ASR is a misleading metric if you’re looking at the wrong or irrelevant classes of errors. Or if you’re comparing a non-streaming ASR model (which is inherently easier to be accurate on) with a streaming ASR model. Unfortunately many benchmarks do not distinguish between ASR streaming vs non-streaming use cases.