A TOPIC, IN CONTEXT

performance benchmarking

Judgments in this source concerning performance benchmarking. Explore 2 viewpoints with evidence from 2 sources.

0 people · 2 sources · 2 viewpoints

Content updated:

Explore connections ↗

Viewpoint map

0 people · 2 sources · 2 viewpoints

Compass improved accuracy by 14–16 points on financial RAG workload

On the High Finance benchmark—a Cohere-built investment-banking RAG workload—Compass improved accuracy by 14–16 points over Azure Search (64.8 → 81.1). That gap can distinguish between unsatisfactory and great end-user answers.

Supporting evidence

Compass is coming to the cloud | Cohere

Original excerpt

The figure below is one instance of the accuracy gain over traditional or standalone search infrastructure on a representative financial-industry RAG workload: embed a query, retrieve presentation materials from an index, and score the top results. On High Finance, a Cohere-built investment-banking benchmark, Compass achieved a 14-16 point improvement on Azure Search (from 64.8 to 81.1). A gap of this size can be the difference between an unsatisfactory answer and a great one for the end user.

Faster batch execution in one crawl, with a measurement caveat

In one 50-action Wikipedia crawl reported by the authors, batch execution took 14,221.7 ms and Playwright took 22,650.3 ms, a 1.59x difference. Both completed all 50 actions. The authors stress that this was one run on one route from one client, rather than a benchmark.

Supporting evidence

Introducing Stagehand v4: The SDK for browser agents.

Original excerpt

We measured one 50-action Wikipedia crawl. Batch finished in 14,221.7ms against Playwright's 22,650.3ms, and both completed 50 of 50 actions. That works out to 3.52 actions per second versus 2.21, a 1.59x difference, with 44.0ms of overhead for the outer batch call.

These findings reflect the available sources, not an exhaustive or current view.