Cartesia | 语音 AI 模型选型指南
如果你关注的是错误或不相关的错误类别,或者将非流式 ASR 模型(其本身更容易做到准确)与流式 ASR 模型进行对比,那么 STT/ASR 的词错误率(WER)就是一个具有误导性的指标。遗憾的是,许多基准测试并未区分 ASR 的流式与非流式使用场景。
原始摘录
Word Error Rate (WER) for STT/ASR is a misleading metric if you’re looking at the wrong or irrelevant classes of errors. Or if you’re comparing a non-streaming ASR model (which is inherently easier to be accurate on) with a streaming ASR model. Unfortunately many benchmarks do not distinguish between ASR streaming vs non-streaming use cases.