Improved token efficiency for longer agent runs · Cursor

Cursor Blog ·

The authors describe token-efficiency work in Cursor’s agent harness. They report a 7% reduction in user token costs without reducing agent quality, a roughly 66% reduction in the system prompt, a 46.9% reduction in total tokens for sessions calling an MCP tool after dynamic tool loading, and a 20% reduction in cold cache misses after cache-related changes. Read 4 viewpoints with supporting evidence and source links.

Jediah Katz, Connor O’Keefe & Calvin Yee

Understand this piece

4 key points

Synthesis

  1. 7% token cost reduction without quality loss

    Improvements to Cursor's agent harness across multiple layers reduced user token costs by 7% while maintaining agent quality.

    Supporting evidence 1

    Original excerpt

    Changes across each of these layers reduced token costs for users by 7% without reducing agent quality.

    Jediah Katz, Connor O’Keefe & Calvin Yee · Paragraph 7

    Context

    Over the past few months we've responded to this shift by improving the efficiency of Cursor's agent harness. The harness gives us direct control over how each request is assembled, how context is reused, and when work is divided across agents.

    Read in source context →
  2. 66% system prompt reduction enabled by model capability gains

    As models improved, explicit prohibitions and prescriptive instructions became unnecessary; Cursor trimmed approximately 66% of its system prompt by relying on model compliance with concise tool behavior definitions.

    Supporting evidence 1

    Original excerpt

    This was true across model families, allowing us to trim roughly 66% of our system prompt.

    Jediah Katz, Connor O’Keefe & Calvin Yee · Paragraph 11

    Context

    As models improved, much of that direction became unnecessary. Instead of long lists of "DO NOT do this," "You must," or "Important" instructions, we could simply define how a tool behaves and models would generally comply.

    Read in source context →

    Continue exploring

    system prompt design →
  3. 46.9% token reduction from dynamic MCP tool loading

    Moving MCP tools into dynamic context—loading them only when invoked—reduced total tokens by 46.9% across sessions that used an MCP tool.

    Supporting evidence 1

    Original excerpt

    This reduced total tokens by 46.9% across sessions that called an MCP tool.

    Jediah Katz, Connor O’Keefe & Calvin Yee · Paragraph 15

    Context

    That created an opportunity to improve efficiency by keeping tools available without including their full definitions in every request. We'd solved a similar problem earlier this year when we moved MCP tools into dynamic context, loading them only when needed.

    Read in source context →

    Continue exploring

    dynamic context loading →
  4. Cache-related changes reduced cold cache misses by 20%

    Cache-related changes reduced cold cache misses by 20%. The authors reserved tools and system instructions for content that rarely changes, then moved more variable setup past the cache boundaries into a “phantom user message.”

    Supporting evidence 1

    Original excerpt

    These changes reduced the rate of cold cache misses by 20%.

    Jediah Katz, Connor O’Keefe & Calvin Yee · Paragraph 35

    Read in source context →

Key passages4

Attributed passages with the context to verify them. Open the original text to check the source.

token efficiency optimization

7% token cost reduction without quality loss

Original excerpt

Changes across each of these layers reduced token costs for users by 7% without reducing agent quality.
Context

Over the past few months we've responded to this shift by improving the efficiency of Cursor's agent harness. The harness gives us direct control over how each request is assembled, how context is reused, and when work is divided across agents.

system prompt design

66% system prompt reduction enabled by model capability gains

Original excerpt

This was true across model families, allowing us to trim roughly 66% of our system prompt.
Context

As models improved, much of that direction became unnecessary. Instead of long lists of "DO NOT do this," "You must," or "Important" instructions, we could simply define how a tool behaves and models would generally comply.

dynamic context loading

46.9% token reduction from dynamic MCP tool loading

Original excerpt

This reduced total tokens by 46.9% across sessions that called an MCP tool.
Context

That created an opportunity to improve efficiency by keeping tools available without including their full definitions in every request. We'd solved a similar problem earlier this year when we moved MCP tools into dynamic context, loading them only when needed.

Source & methodology

These viewpoints are linked to their original sources. Paraphrases are labeled and are not verbatim quotes.

Open transcript or source material (opens in a new tab)Report an issue