提供诸如推测解码接受长度和每个副本 token 延迟分位数等引擎级指标
调试推理问题真正需要的指标——例如推测解码接受长度以及每个副本的引擎侧 token 延迟分位数——会在仪表板中自动提供。
支持这项说法
Modal Auto Endpoints 介绍:真正属于你的优化推理 | Modal 博客
我们不会隐藏指标。你调试推理问题真正需要的指标,比如推测解码接受长度,以及每个副本、引擎侧的 token 延迟分位数,都会自动显示在仪表盘中。门槛很低,但这可不是我们设的!
原始摘录
We don't hide the metrics . The metrics you actually need to debug inference issues, like speculative decoding acceptance length and per-replica, engine-side token latency quantiles, are automatically provided in a dashboard. Low bar, but we didn't put it there!