A TOPIC, IN CONTEXT

evaluation methodology

Judgments in this source concerning evaluation methodology. Explore 2 viewpoints with evidence from 2 sources.

0 people · 2 sources · 2 viewpoints

Content updated:

Explore connections ↗

Viewpoint map

0 people · 2 sources · 2 viewpoints

Track outcomes before routing with evals or A/B tests

Before deploying routing, put measures of success in place—such as offline evaluations, online evaluators, or user feedback logged on traces—and if building an eval dataset is too costly, an A/B test on live traffic works well.

Supporting evidence

How to Build a Model Router in the Harness

Original excerpt

Track task outcomes. Put measures of success in place before you route: evals , online evaluators , or user feedback on traces . If building an eval dataset is too costly or difficult, an A/B test on live traffic works well.

LLM judges score redline quality against lawyer-defined rubrics

Redline quality is evaluated using lawyer-sourced rubrics that assess minimality, correct placement, and legal soundness. A committee of three frontier LLMs acts as independent judges to score outputs; their votes are aggregated.

Supporting evidence

Rebuilding Playbook Review as a Multi-Agent System | Harvey

Original excerpt

Redline quality is more complex in that it requires judgment on whether an edit is minimal, correctly placed, and legally sound . After sourcing the rubrics with lawyers, we set up an evaluation framework that uses LLM judges to score against these rubrics. We then used a committee of three frontier models to score independently and aggregate the votes.
Context

Flagging risk is best understood and scored as a classification problem which is a straightforward task.

These findings reflect the available sources, not an exhaustive or current view.