How to Build a Model Router in the Harness

LangChain Blog ·

The authors explain why they place model routing in the agent harness, where domain and task context is available. They report their Open SWE cost comparison, describe comparing model intelligence with task cost, and recommend tracking task outcomes before routing through evaluations, user feedback or an A/B test. Lee 4 puntos de vista con sus evidencias y enlaces a las fuentes.

Sydney Runkle, Eugene Yurtsev

De un vistazo

  • Routing decision belongs in the agent harness

    The routing decision belongs in the agent harness—not a generic gateway—because choosing the right model requires domain and task context that the harness already assembles and that a gateway typically lacks.

    Ver el momento de apoyo · Párrafo 2
  • 64% lower median cost, with no measurable quality change

    The authors report that a model router for LangChain’s Open SWE coding agent reduced median cost per thread by 64% against their previous baseline of always using a top-tier frontier model. They report no measurable change in quality in that comparison.

    Ver el momento de apoyo · Párrafo 3
  • Use Pareto frontier analysis for model selection

    The Artificial Analysis Intelligence Index scores models on common tasks and reports cost per task, enabling plotting on an intelligence-vs-cost curve; the Pareto frontier identifies models that are simultaneously the cheapest and smartest.

    Ver el momento de apoyo · Párrafo 12
  • Track outcomes before routing with evals or A/B tests

    Before deploying routing, put measures of success in place—such as offline evaluations, online evaluators, or user feedback logged on traces—and if building an eval dataset is too costly, an A/B test on live traffic works well.

    Ver el momento de apoyo · Párrafo 52

Pasajes clave4

Pasajes atribuidos con contexto para verificarlos. Abra el texto original para comprobar la fuente.

model evaluation

Use Pareto frontier analysis for model selection

Extracto original

The Artificial Analysis Intelligence Index scores models on a common set of tasks and reports the cost per task, so you can plot them all on one curve of intelligence against cost. The Pareto frontier is the set of models that are the cheapest and smartest.
model routing architecture

Routing decision belongs in the agent harness

Extracto original

We believe that routing decision belongs in the agent harness , not a generic gateway, because choosing the right model requires the same domain and task context the harness already assembles and that a gateway typically lacks.
Contexto

Past a certain point, you hit diminishing returns: a more capable model adds little quality while cost and latency keep climbing. A good agent has model-harness-task fit : the right model with the right context for a given task. A model router picks that model for each task.

cost optimization

64% lower median cost, with no measurable quality change

Extracto original

Compared to our previous baseline of always using a top-tier frontier model, it cut median cost per thread by 64% with no measurable change in quality.
Contexto

We felt this pain recently at LangChain as our monthly coding agent spend started to climb rapidly. Hearing the same concern from customers, we set out to build an effective model router for Open SWE , our open source coding agent. This post covers how we built the router, what we learned, and how you can get started building model routing into your agents.

Fuente y metodología

Estas perspectivas enlazan a sus fuentes originales. Las paráfrasis están identificadas y no son citas textuales.

Abrir transcripción o material de origen (se abre en una pestaña nueva)Reportar un problema