如何在 Harness 中构建模型路由器

LangChain Blog ·

作者解释了为何将模型路由置于 agent harness 中,因为该处可获得领域与任务上下文。他们报告了 Open SWE 的成本对比结果,描述了如何将模型智能与任务成本进行比较,并建议在通过评估、用户反馈或 A/B 测试进行路由之前先跟踪任务结果。 阅读 4 条观点,查看支持证据与原始来源。

Sydney Runkle, Eugene Yurtsev

理解这篇

4 个要点

综合解读

  1. 路由决策属于 agent harness

    路由决策应属于 agent harness,而非通用网关,因为选择合适的模型需要领域和任务上下文,而 harness 已经汇集了这些上下文,网关通常则不具备。

    支持这项说法 1

    我们认为路由决策应属于 agent harness,而非通用网关,因为选择合适的模型需要 harness 已经汇集而网关通常缺乏的相同领域和任务上下文。

    Sydney Runkle, Eugene Yurtsev · 段落 2

    原始摘录
    We believe that routing decision belongs in the agent harness , not a generic gateway, because choosing the right model requires the same domain and task context the harness already assembles and that a gateway typically lacks.
    上下文

    超过某个临界点后,你会遇到收益递减:能力更强的模型几乎不会提升质量,而成本和延迟却持续攀升。优秀的 agent 具备模型-harness-任务契合度:针对特定任务使用合适的模型与合适的上下文。模型路由器会为每项任务挑选该模型。

    原始上下文

    Past a certain point, you hit diminishing returns: a more capable model adds little quality while cost and latency keep climbing. A good agent has model-harness-task fit : the right model with the right context for a given task. A model router picks that model for each task.

    回到原文语境 →
  2. 中位成本降低 64%,质量无可衡量变化

    作者报告称,与此前始终使用顶级前沿模型的基线相比,为 LangChain 的 Open SWE 编码 agent 配备的模型路由器将每个线程的中位成本降低了 64%。他们报告在该比较中质量没有可衡量的变化。

    支持这项说法 1

    与我们此前始终使用顶级前沿模型的基线相比,它将每个线程的中位成本降低了 64%,且质量没有可衡量的变化。

    Sydney Runkle, Eugene Yurtsev · 段落 3

    原始摘录
    Compared to our previous baseline of always using a top-tier frontier model, it cut median cost per thread by 64% with no measurable change in quality.
    上下文

    我们最近在 LangChain 感受到了这种痛苦,因为我们每月在编码 agent 上的支出开始迅速攀升。听到客户也有同样的担忧后,我们着手为开源编码 agent Open SWE 构建一个有效的模型路由器。本文介绍了我们如何构建该路由器、学到了什么,以及你如何开始在 agent 中构建模型路由。

    原始上下文

    We felt this pain recently at LangChain as our monthly coding agent spend started to climb rapidly. Hearing the same concern from customers, we set out to build an effective model router for Open SWE , our open source coding agent. This post covers how we built the router, what we learned, and how you can get started building model routing into your agents.

    回到原文语境 →

    继续探索

    成本优化 →
  3. 使用帕累托前沿分析进行模型选择

    Artificial Analysis Intelligence Index 对模型在常见任务上进行评分并报告每项任务的成本,从而可在智能与成本曲线上绘图;帕累托前沿可识别出同时最便宜且最智能的模型。

    支持这项说法 1

    人工分析智能指数(Artificial Analysis Intelligence Index)基于一组通用任务对模型进行评分,并报告每项任务的运行成本,因此你可以将所有模型绘制在一条‘智能—成本’关系曲线上。帕累托前沿指其中成本最低、智能表现最优的模型集合。

    Sydney Runkle, Eugene Yurtsev · 段落 12

    原始摘录
    The Artificial Analysis Intelligence Index scores models on a common set of tasks and reports the cost per task, so you can plot them all on one curve of intelligence against cost. The Pareto frontier is the set of models that are the cheapest and smartest.
    回到原文语境 →

    继续探索

    模型评估 →
  4. 在路由前通过评估或 A/B 测试跟踪结果

    在部署路由之前,应先建立成功衡量标准——例如离线评估、在线评估器,或在 trace 上记录的用户反馈——如果构建评估数据集成本过高,对线上流量进行 A/B 测试也很有效。

    支持这项说法 1

    跟踪任务结果。应在路由前就明确成功度量标准,例如评估(evals)、在线评估员(online evaluators)或用户对追踪记录(traces)的反馈。若构建评估数据集成本过高或难度过大,可在真实流量上开展 A/B 测试,效果通常较好。

    Sydney Runkle, Eugene Yurtsev · 段落 52

    原始摘录
    Track task outcomes. Put measures of success in place before you route: evals , online evaluators , or user feedback on traces . If building an eval dataset is too costly or difficult, an A/B test on live traffic works well.
    回到原文语境 →

    继续探索

    评估方法 →

关键段落4

带明确归属与语境的原文片段。打开原始文本核查出处。

模型评估

使用帕累托前沿分析进行模型选择

人工分析智能指数(Artificial Analysis Intelligence Index)基于一组通用任务对模型进行评分,并报告每项任务的运行成本,因此你可以将所有模型绘制在一条‘智能—成本’关系曲线上。帕累托前沿指其中成本最低、智能表现最优的模型集合。

原始摘录
The Artificial Analysis Intelligence Index scores models on a common set of tasks and reports the cost per task, so you can plot them all on one curve of intelligence against cost. The Pareto frontier is the set of models that are the cheapest and smartest.
评估方法

在路由前通过评估或 A/B 测试跟踪结果

跟踪任务结果。应在路由前就明确成功度量标准,例如评估(evals)、在线评估员(online evaluators)或用户对追踪记录(traces)的反馈。若构建评估数据集成本过高或难度过大,可在真实流量上开展 A/B 测试,效果通常较好。

原始摘录
Track task outcomes. Put measures of success in place before you route: evals , online evaluators , or user feedback on traces . If building an eval dataset is too costly or difficult, an A/B test on live traffic works well.
模型路由架构

路由决策属于 agent harness

我们认为路由决策应属于 agent harness,而非通用网关,因为选择合适的模型需要 harness 已经汇集而网关通常缺乏的相同领域和任务上下文。

原始摘录
We believe that routing decision belongs in the agent harness , not a generic gateway, because choosing the right model requires the same domain and task context the harness already assembles and that a gateway typically lacks.
上下文

超过某个临界点后,你会遇到收益递减:能力更强的模型几乎不会提升质量,而成本和延迟却持续攀升。优秀的 agent 具备模型-harness-任务契合度:针对特定任务使用合适的模型与合适的上下文。模型路由器会为每项任务挑选该模型。

原始上下文

Past a certain point, you hit diminishing returns: a more capable model adds little quality while cost and latency keep climbing. A good agent has model-harness-task fit : the right model with the right context for a given task. A model router picks that model for each task.

成本优化

中位成本降低 64%,质量无可衡量变化

与我们此前始终使用顶级前沿模型的基线相比,它将每个线程的中位成本降低了 64%,且质量没有可衡量的变化。

原始摘录
Compared to our previous baseline of always using a top-tier frontier model, it cut median cost per thread by 64% with no measurable change in quality.
上下文

我们最近在 LangChain 感受到了这种痛苦,因为我们每月在编码 agent 上的支出开始迅速攀升。听到客户也有同样的担忧后,我们着手为开源编码 agent Open SWE 构建一个有效的模型路由器。本文介绍了我们如何构建该路由器、学到了什么,以及你如何开始在 agent 中构建模型路由。

原始上下文

We felt this pain recently at LangChain as our monthly coding agent spend started to climb rapidly. Hearing the same concern from customers, we set out to build an effective model router for Open SWE , our open source coding agent. This post covers how we built the router, what we learned, and how you can get started building model routing into your agents.

来源与研究方法

这些观点均关联原始来源。转述已明确标注,不作为逐字原话展示。

打开转录或来源材料 (在新标签页中打开)报告问题

继续研究相关话题

以下资料涉及本文的话题;讨论相同话题不代表观点一致。