Improved token efficiency for longer agent runs · Cursor

Cursor Blog ·

The authors describe token-efficiency work in Cursor’s agent harness. They report a 7% reduction in user token costs without reducing agent quality, a roughly 66% reduction in the system prompt, a 46.9% reduction in total tokens for sessions calling an MCP tool after dynamic tool loading, and a 20% reduction in cold cache misses after cache-related changes. Lisez 4 points de vue avec leurs éléments à l’appui et les liens vers les sources.

Jediah Katz, Connor O’Keefe & Calvin Yee

En un coup d’œil

  • 7% token cost reduction without quality loss

    Improvements to Cursor's agent harness across multiple layers reduced user token costs by 7% while maintaining agent quality.

    Lire le moment probant · Paragraphe 7
  • 66% system prompt reduction enabled by model capability gains

    As models improved, explicit prohibitions and prescriptive instructions became unnecessary; Cursor trimmed approximately 66% of its system prompt by relying on model compliance with concise tool behavior definitions.

    Lire le moment probant · Paragraphe 11
  • 46.9% token reduction from dynamic MCP tool loading

    Moving MCP tools into dynamic context—loading them only when invoked—reduced total tokens by 46.9% across sessions that used an MCP tool.

    Lire le moment probant · Paragraphe 15
  • Cache-related changes reduced cold cache misses by 20%

    Cache-related changes reduced cold cache misses by 20%. The authors reserved tools and system instructions for content that rarely changes, then moved more variable setup past the cache boundaries into a “phantom user message.”

    Lire le moment probant · Paragraphe 35

Passages clés4

Passages attribués et accompagnés du contexte nécessaire à leur vérification. Ouvrez le texte original pour vérifier la source.

token efficiency optimization

7% token cost reduction without quality loss

Extrait original

Changes across each of these layers reduced token costs for users by 7% without reducing agent quality.
Contexte

Over the past few months we've responded to this shift by improving the efficiency of Cursor's agent harness. The harness gives us direct control over how each request is assembled, how context is reused, and when work is divided across agents.

system prompt design

66% system prompt reduction enabled by model capability gains

Extrait original

This was true across model families, allowing us to trim roughly 66% of our system prompt.
Contexte

As models improved, much of that direction became unnecessary. Instead of long lists of "DO NOT do this," "You must," or "Important" instructions, we could simply define how a tool behaves and models would generally comply.

dynamic context loading

46.9% token reduction from dynamic MCP tool loading

Extrait original

This reduced total tokens by 46.9% across sessions that called an MCP tool.
Contexte

That created an opportunity to improve efficiency by keeping tools available without including their full definitions in every request. We'd solved a similar problem earlier this year when we moved MCP tools into dynamic context, loading them only when needed.

Source et méthodologie

Ces points de vue renvoient à leurs sources originales. Les reformulations sont signalées et ne sont pas des citations mot à mot.

Ouvrir la transcription ou les documents sources (s’ouvre dans un nouvel onglet)Signaler un problème