Hyper-𝜏-bench: Evaluating agents that build agents

Sierra Blog ·

The article contrasts evaluating customer-service agents with evaluating models that build agents. It reports standalone and engineer-paired task performance, and relates client questioning to success on benchmark tasks. Lisez 3 points de vue avec leurs éléments à l’appui et les liens vers les sources.

Ben Shi, Keshav Dhandhania

En un coup d’œil

  • From customer service agent to agent builder

    The field has moved beyond evaluating whether models can act as reliable customer service agents (𝜏-bench's original focus) to evaluating whether models can themselves build such agents.

    Lire le moment probant · Paragraphe 1
  • Reported success with engineer context

    The authors report 23.9% success on held-out tasks for their best standalone configuration, Claude Opus 5 with max reasoning in Claude Code. Paired with an engineer with deep context, the same class of model reaches 82.2% on those tasks.

    Lire le moment probant · Paragraphe 6
  • Client questioning strongly correlates with agent-building success

    On tasks where the simulated client held sole context for 20–25 requirements, developer agents that asked zero questions scored 5%, one question 15%, two questions 25%, and so on — demonstrating a direct payoff from client interviews.

    Lire le moment probant · Paragraphe 9

Passages clés3

Passages attribués et accompagnés du contexte nécessaire à leur vérification. Ouvrez le texte original pour vérifier la source.

agent evaluation scope

From customer service agent to agent builder

Extrait original

We built 𝜏-bench in 2024 to answer a question that felt novel at the time: Can a model act as a reliable customer service agent? That’s table stakes now. The harder question is, who’s building the agent in the first place? Increasingly, it’s the models themselves.
client interviewing behavior

Client questioning strongly correlates with agent-building success

Extrait original

They don’t interview the client. Developers only asked four questions at most for tasks where the client had sole context for 20-25 requirements. Asking pays off directly. For tasks where reference agents (built by an engineer) scored 95–100% — the builds that asked zero questions scored 5%, one question 15%, two questions 25%, and so on.
human-AI collaboration performance

Reported success with engineer context

Extrait original

Working alone, our best configuration — Claude Opus 5 (max reasoning) running in Claude Code — passes just 23.9% of the held-out evaluation tasks. Paired with an engineer with deep context, the same class of model reaches 82.2% on the same tasks.

Source et méthodologie

Ces points de vue renvoient à leurs sources originales. Les reformulations sont signalées et ne sont pas des citations mot à mot.

Ouvrir la transcription ou les documents sources (s’ouvre dans un nouvel onglet)Signaler un problème