A TOPIC, IN CONTEXT

cost optimization

Judgments in this source concerning cost optimization. Explore 2 viewpoints with evidence from 2 sources.

0 people · 2 sources · 2 viewpoints

Content updated:

Explore connections ↗

Viewpoint map

0 people · 2 sources · 2 viewpoints

64% lower median cost, with no measurable quality change

The authors report that a model router for LangChain’s Open SWE coding agent reduced median cost per thread by 64% against their previous baseline of always using a top-tier frontier model. They report no measurable change in quality in that comparison.

Supporting evidence

How to Build a Model Router in the Harness

Original excerpt

Compared to our previous baseline of always using a top-tier frontier model, it cut median cost per thread by 64% with no measurable change in quality.
Context

We felt this pain recently at LangChain as our monthly coding agent spend started to climb rapidly. Hearing the same concern from customers, we set out to build an effective model router for Open SWE , our open source coding agent. This post covers how we built the router, what we learned, and how you can get started building model routing into your agents.

Precise retrieval reduces token costs

More precise retrieval delivers smaller, higher-quality inputs to generative models, cutting inference costs and task completion time. Retrieval is among the most effective cost levers available to businesses today.

Supporting evidence

Compass is coming to the cloud | Cohere

Original excerpt

Token economics: Every irrelevant result passed to a model consumes tokens and occupies limited context space. More precise retrieval creates smaller, higher-quality inputs, reducing inference costs and cutting the time needed to complete a task. Retrieval is one of the most effective cost levers available to businesses today.

These findings reflect the available sources, not an exhaustive or current view.