How to Build a Model Router in the Harness

LangChain Blog ·

The authors explain why they place model routing in the agent harness, where domain and task context is available. They report their Open SWE cost comparison, describe comparing model intelligence with task cost, and recommend tracking task outcomes before routing through evaluations, user feedback or an A/B test. Read 4 viewpoints with supporting evidence and source links.

Sydney Runkle, Eugene Yurtsev

Understand this piece

4 key points

Synthesis

  1. Routing decision belongs in the agent harness

    The routing decision belongs in the agent harness—not a generic gateway—because choosing the right model requires domain and task context that the harness already assembles and that a gateway typically lacks.

    Supporting evidence 1

    Original excerpt

    We believe that routing decision belongs in the agent harness , not a generic gateway, because choosing the right model requires the same domain and task context the harness already assembles and that a gateway typically lacks.

    Sydney Runkle, Eugene Yurtsev · Paragraph 2

    Context

    Past a certain point, you hit diminishing returns: a more capable model adds little quality while cost and latency keep climbing. A good agent has model-harness-task fit : the right model with the right context for a given task. A model router picks that model for each task.

    Read in source context →
  2. 64% lower median cost, with no measurable quality change

    The authors report that a model router for LangChain’s Open SWE coding agent reduced median cost per thread by 64% against their previous baseline of always using a top-tier frontier model. They report no measurable change in quality in that comparison.

    Supporting evidence 1

    Original excerpt

    Compared to our previous baseline of always using a top-tier frontier model, it cut median cost per thread by 64% with no measurable change in quality.

    Sydney Runkle, Eugene Yurtsev · Paragraph 3

    Context

    We felt this pain recently at LangChain as our monthly coding agent spend started to climb rapidly. Hearing the same concern from customers, we set out to build an effective model router for Open SWE , our open source coding agent. This post covers how we built the router, what we learned, and how you can get started building model routing into your agents.

    Read in source context →

    Continue exploring

    cost optimization →
  3. Use Pareto frontier analysis for model selection

    The Artificial Analysis Intelligence Index scores models on common tasks and reports cost per task, enabling plotting on an intelligence-vs-cost curve; the Pareto frontier identifies models that are simultaneously the cheapest and smartest.

    Supporting evidence 1

    Original excerpt

    The Artificial Analysis Intelligence Index scores models on a common set of tasks and reports the cost per task, so you can plot them all on one curve of intelligence against cost. The Pareto frontier is the set of models that are the cheapest and smartest.

    Sydney Runkle, Eugene Yurtsev · Paragraph 12

    Read in source context →

    Continue exploring

    model evaluation →
  4. Track outcomes before routing with evals or A/B tests

    Before deploying routing, put measures of success in place—such as offline evaluations, online evaluators, or user feedback logged on traces—and if building an eval dataset is too costly, an A/B test on live traffic works well.

    Supporting evidence 1

    Original excerpt

    Track task outcomes. Put measures of success in place before you route: evals , online evaluators , or user feedback on traces . If building an eval dataset is too costly or difficult, an A/B test on live traffic works well.

    Sydney Runkle, Eugene Yurtsev · Paragraph 52

    Read in source context →

    Continue exploring

    evaluation methodology →

Key passages4

Attributed passages with the context to verify them. Open the original text to check the source.

model evaluation

Use Pareto frontier analysis for model selection

Original excerpt

The Artificial Analysis Intelligence Index scores models on a common set of tasks and reports the cost per task, so you can plot them all on one curve of intelligence against cost. The Pareto frontier is the set of models that are the cheapest and smartest.
evaluation methodology

Track outcomes before routing with evals or A/B tests

Original excerpt

Track task outcomes. Put measures of success in place before you route: evals , online evaluators , or user feedback on traces . If building an eval dataset is too costly or difficult, an A/B test on live traffic works well.
model routing architecture

Routing decision belongs in the agent harness

Original excerpt

We believe that routing decision belongs in the agent harness , not a generic gateway, because choosing the right model requires the same domain and task context the harness already assembles and that a gateway typically lacks.
Context

Past a certain point, you hit diminishing returns: a more capable model adds little quality while cost and latency keep climbing. A good agent has model-harness-task fit : the right model with the right context for a given task. A model router picks that model for each task.

cost optimization

64% lower median cost, with no measurable quality change

Original excerpt

Compared to our previous baseline of always using a top-tier frontier model, it cut median cost per thread by 64% with no measurable change in quality.
Context

We felt this pain recently at LangChain as our monthly coding agent spend started to climb rapidly. Hearing the same concern from customers, we set out to build an effective model router for Open SWE , our open source coding agent. This post covers how we built the router, what we learned, and how you can get started building model routing into your agents.

Source & methodology

These viewpoints are linked to their original sources. Paraphrases are labeled and are not verbatim quotes.

Open transcript or source material (opens in a new tab)Report an issue

Continue with this topic

More sources on topics discussed here. Shared topics do not imply agreement.