为更长的 agent 运行提升 token 效率 · Cursor

Cursor Blog ·

作者描述了 Cursor agent 框架中的 token 效率优化工作。他们报告称,在不降低 agent 质量的前提下,用户 token 成本降低了 7%,系统提示词缩减了约 66%,在动态工具加载后调用 MCP 工具的会话中总 token 数减少了 46.9%,并在与缓存相关的改动后冷缓存未命中减少了 20%。 阅读 4 条观点,查看支持证据与原始来源。

Jediah Katz, Connor O’Keefe & Calvin Yee

理解这篇

4 个要点

综合解读

  1. token 成本降低 7% 且无质量损失

    对 Cursor agent 框架多个层面的改进使用户 token 成本降低了 7%,同时保持了 agent 质量。

    支持这项说法 1

    对上述各个层面的改动使用户的 token 成本降低了 7%,且未降低 agent 质量。

    Jediah Katz, Connor O’Keefe & Calvin Yee · 段落 7

    原始摘录
    Changes across each of these layers reduced token costs for users by 7% without reducing agent quality.
    上下文

    在过去几个月里,我们通过提升 Cursor agent 框架的效率来应对这一转变。该框架让我们能够直接控制每个请求的组装方式、上下文的复用方式,以及何时在多个 agent 之间分配工作。

    原始上下文

    Over the past few months we've responded to this shift by improving the efficiency of Cursor's agent harness. The harness gives us direct control over how each request is assembled, how context is reused, and when work is divided across agents.

    回到原文语境 →
  2. 模型能力提升使系统提示词缩减 66%

    随着模型能力的提升,显式的禁止性规定和指令性说明变得不再必要;Cursor 依靠模型遵循简洁的工具行为定义,将其系统提示词删减了约 66%。

    支持这项说法 1

    这一现象在各类模型家族中均成立,系统提示词因此减少了约66%。

    Jediah Katz, Connor O’Keefe & Calvin Yee · 段落 11

    原始摘录
    This was true across model families, allowing us to trim roughly 66% of our system prompt.
    上下文

    随着模型能力提升,大部分指令性内容已不再必要:我们无需再写冗长的‘切勿执行此操作’‘必须执行’或‘重要’等说明,只需清晰定义工具行为,模型通常即可正确执行。

    原始上下文

    As models improved, much of that direction became unnecessary. Instead of long lists of "DO NOT do this," "You must," or "Important" instructions, we could simply define how a tool behaves and models would generally comply.

    回到原文语境 →
  3. 动态加载 MCP 工具使 token 减少 46.9%

    将 MCP 工具移入动态上下文(仅在调用时加载),使使用了 MCP 工具的会话的总 token 数减少了 46.9%。

    支持这项说法 1

    在调用了 MCP 工具的各会话中,此举使总 token 数减少了 46.9%。

    Jediah Katz, Connor O’Keefe & Calvin Yee · 段落 15

    原始摘录
    This reduced total tokens by 46.9% across sessions that called an MCP tool.
    上下文

    这创造了一个提升效率的机会:让工具保持可用状态,而无需在每次请求中都包含其完整定义。今年早些时候,当我们把 MCP 工具移入动态上下文、仅在需要时才加载它们时,就解决了类似的问题。

    原始上下文

    That created an opportunity to improve efficiency by keeping tools available without including their full definitions in every request. We'd solved a similar problem earlier this year when we moved MCP tools into dynamic context, loading them only when needed.

    回到原文语境 →
  4. 与缓存相关的改动使冷缓存未命中减少 20%

    与缓存相关的改动使冷缓存未命中减少了 20%。作者将工具和系统指令保留给极少变化的内容,然后将更多易变的设置移至缓存边界之外,放入一条“phantom user message”中。

    支持这项说法 1

    这些变更使冷缓存未命中率下降了20%。

    Jediah Katz, Connor O’Keefe & Calvin Yee · 段落 35

    原始摘录
    These changes reduced the rate of cold cache misses by 20%.
    回到原文语境 →

关键段落4

带明确归属与语境的原文片段。打开原始文本核查出处。

token 效率优化

token 成本降低 7% 且无质量损失

对上述各个层面的改动使用户的 token 成本降低了 7%,且未降低 agent 质量。

原始摘录
Changes across each of these layers reduced token costs for users by 7% without reducing agent quality.
上下文

在过去几个月里,我们通过提升 Cursor agent 框架的效率来应对这一转变。该框架让我们能够直接控制每个请求的组装方式、上下文的复用方式,以及何时在多个 agent 之间分配工作。

原始上下文

Over the past few months we've responded to this shift by improving the efficiency of Cursor's agent harness. The harness gives us direct control over how each request is assembled, how context is reused, and when work is divided across agents.

系统提示词设计

模型能力提升使系统提示词缩减 66%

这一现象在各类模型家族中均成立,系统提示词因此减少了约66%。

原始摘录
This was true across model families, allowing us to trim roughly 66% of our system prompt.
上下文

随着模型能力提升,大部分指令性内容已不再必要:我们无需再写冗长的‘切勿执行此操作’‘必须执行’或‘重要’等说明,只需清晰定义工具行为,模型通常即可正确执行。

原始上下文

As models improved, much of that direction became unnecessary. Instead of long lists of "DO NOT do this," "You must," or "Important" instructions, we could simply define how a tool behaves and models would generally comply.

动态上下文加载

动态加载 MCP 工具使 token 减少 46.9%

在调用了 MCP 工具的各会话中,此举使总 token 数减少了 46.9%。

原始摘录
This reduced total tokens by 46.9% across sessions that called an MCP tool.
上下文

这创造了一个提升效率的机会:让工具保持可用状态,而无需在每次请求中都包含其完整定义。今年早些时候,当我们把 MCP 工具移入动态上下文、仅在需要时才加载它们时,就解决了类似的问题。

原始上下文

That created an opportunity to improve efficiency by keeping tools available without including their full definitions in every request. We'd solved a similar problem earlier this year when we moved MCP tools into dynamic context, loading them only when needed.

来源与研究方法

这些观点均关联原始来源。转述已明确标注,不作为逐字原话展示。

打开转录或来源材料 (在新标签页中打开)报告问题