UN TEMA, EN CONTEXTO

evaluation methodology

Judgments in this source concerning evaluation methodology. Explora 2 puntos de vista con evidencias de 2 fuentes.

0 personas · 2 fuentes · 2 opiniones expresadas

Contenido actualizado:

Explorar conexiones ↗

Mapa de perspectivas

0 personas · 2 fuentes · 2 opiniones expresadas

Track outcomes before routing with evals or A/B tests

Before deploying routing, put measures of success in place—such as offline evaluations, online evaluators, or user feedback logged on traces—and if building an eval dataset is too costly, an A/B test on live traffic works well.

Evidencia a favor

How to Build a Model Router in the Harness

Extracto original

Track task outcomes. Put measures of success in place before you route: evals , online evaluators , or user feedback on traces . If building an eval dataset is too costly or difficult, an A/B test on live traffic works well.

LLM judges score redline quality against lawyer-defined rubrics

Redline quality is evaluated using lawyer-sourced rubrics that assess minimality, correct placement, and legal soundness. A committee of three frontier LLMs acts as independent judges to score outputs; their votes are aggregated.

Evidencia a favor

Rebuilding Playbook Review as a Multi-Agent System | Harvey

Extracto original

Redline quality is more complex in that it requires judgment on whether an edit is minimal, correctly placed, and legally sound . After sourcing the rubrics with lawyers, we set up an evaluation framework that uses LLM judges to score against these rubrics. We then used a committee of three frontier models to score independently and aggregate the votes.
Contexto

Flagging risk is best understood and scored as a classification problem which is a straightforward task.

Estos resultados reflejan las fuentes disponibles, no una visión exhaustiva ni necesariamente actual.