UN THÈME, DANS SON CONTEXTE

evaluation methodology

Judgments in this source concerning evaluation methodology. Explorez 2 points de vue avec des éléments tirés de 2 sources.

0 personnes · 2 sources · 2 opinions exprimées

Contenu mis à jour:

Explorer les liens ↗

Carte des points de vue

0 personnes · 2 sources · 2 opinions exprimées

Track outcomes before routing with evals or A/B tests

Before deploying routing, put measures of success in place—such as offline evaluations, online evaluators, or user feedback logged on traces—and if building an eval dataset is too costly, an A/B test on live traffic works well.

Éléments favorables

How to Build a Model Router in the Harness

Extrait original

Track task outcomes. Put measures of success in place before you route: evals , online evaluators , or user feedback on traces . If building an eval dataset is too costly or difficult, an A/B test on live traffic works well.

LLM judges score redline quality against lawyer-defined rubrics

Redline quality is evaluated using lawyer-sourced rubrics that assess minimality, correct placement, and legal soundness. A committee of three frontier LLMs acts as independent judges to score outputs; their votes are aggregated.

Éléments favorables

Rebuilding Playbook Review as a Multi-Agent System | Harvey

Extrait original

Redline quality is more complex in that it requires judgment on whether an edit is minimal, correctly placed, and legally sound . After sourcing the rubrics with lawyers, we set up an evaluation framework that uses LLM judges to score against these rubrics. We then used a committee of three frontier models to score independently and aggregate the votes.
Contexte

Flagging risk is best understood and scored as a classification problem which is a straightforward task.

Ces résultats reflètent les sources disponibles, sans constituer une vue exhaustive ou à jour.