话题与观点

性能基准测试

本来源中关于性能基准测试的判断。 阅读 2 条观点,核对 2 个来源中的证据。

0 位人物 · 2 个来源 · 2 条观点

内容更新于:

探索知识关联 ↗

话题观点地图

0 位人物 · 2 个来源 · 2 条观点

Compass 在金融 RAG 工作负载上的准确率提高了 14–16 个百分点

在 High Finance 基准测试(一项由 Cohere 构建的投资银行 RAG 工作负载)中,Compass 的准确率比 Azure Search 提高了 14–16 个百分点(从 64.8 提升至 81.1)。这一差距可能决定了最终用户得到的是不令人满意的答案还是出色的答案。

支持这项说法

Compass 即将登陆云端 | Cohere

下图展示了一个典型金融行业RAG工作负载中,相较于传统或独立搜索基础设施所实现的准确率提升:即对查询进行嵌入(embed),从索引中检索演示材料,并对顶部结果进行评分。在“高阶金融”(High Finance)——一个由Cohere构建的投资银行业务基准测试中,Compass在Azure Search上的表现提升了14至16个百分点(从64.8提升至81.1)。如此幅度的差距,可能直接决定终端用户获得的是令人不满的答案,还是出色的答案。

原始摘录
The figure below is one instance of the accuracy gain over traditional or standalone search infrastructure on a representative financial-industry RAG workload: embed a query, retrieve presentation materials from an index, and score the top results. On High Finance, a Cohere-built investment-banking benchmark, Compass achieved a 14-16 point improvement on Azure Search (from 64.8 to 81.1). A gap of this size can be the difference between an unsatisfactory answer and a great one for the end user.

单次爬取中批处理更快,但数据不具备基准测试效力

作者在一次针对 Wikipedia 的爬取实验中执行了 50 个操作:批处理执行耗时 14,221.7 毫秒,Playwright 耗时 22,650.3 毫秒,前者快 1.59 倍;两者均成功完成全部 50 个操作。作者强调,该数据仅来自单一客户端对单一路径的一次运行,不构成正式基准测试。

支持这项说法

Stagehand v4 介绍:面向浏览器代理的 SDK。

我们测量了一次包含 50 项操作的维基百科爬取。批处理在 14,221.7ms 内完成,而 Playwright 为 22,650.3ms,两者均完成了 50 项操作中的全部 50 项。换算下来即每秒 3.52 项操作对比 2.21 项,相差 1.59 倍,外层批处理调用的开销为 44.0ms。

原始摘录
We measured one 50-action Wikipedia crawl. Batch finished in 14,221.7ms against Playwright's 22,650.3ms, and both completed 50 of 50 actions. That works out to 3.52 actions per second versus 2.21, a 1.59x difference, with 44.0ms of overhead for the outer batch call.

这些结论只反映当前可用来源,不代表全面或最新的观点。