Hyper-𝜏-bench: Evaluating agents that build agents

Sierra Blog ·

The article contrasts evaluating customer-service agents with evaluating models that build agents. It reports standalone and engineer-paired task performance, and relates client questioning to success on benchmark tasks. Lies 3 Standpunkte mit Belegen und Links zu den Originalquellen.

Ben Shi, Keshav Dhandhania

Auf einen Blick

  • From customer service agent to agent builder

    The field has moved beyond evaluating whether models can act as reliable customer service agents (𝜏-bench's original focus) to evaluating whether models can themselves build such agents.

    Unterstützendes Moment lesen · Absatz 1
  • Reported success with engineer context

    The authors report 23.9% success on held-out tasks for their best standalone configuration, Claude Opus 5 with max reasoning in Claude Code. Paired with an engineer with deep context, the same class of model reaches 82.2% on those tasks.

    Unterstützendes Moment lesen · Absatz 6
  • Client questioning strongly correlates with agent-building success

    On tasks where the simulated client held sole context for 20–25 requirements, developer agents that asked zero questions scored 5%, one question 15%, two questions 25%, and so on — demonstrating a direct payoff from client interviews.

    Unterstützendes Moment lesen · Absatz 9

Wichtige Passagen3

Zugeordnete Passagen mit dem Kontext zur Überprüfung. Öffnen Sie den Originaltext, um die Quelle zu prüfen.

agent evaluation scope

From customer service agent to agent builder

Originalauszug

We built 𝜏-bench in 2024 to answer a question that felt novel at the time: Can a model act as a reliable customer service agent? That’s table stakes now. The harder question is, who’s building the agent in the first place? Increasingly, it’s the models themselves.
client interviewing behavior

Client questioning strongly correlates with agent-building success

Originalauszug

They don’t interview the client. Developers only asked four questions at most for tasks where the client had sole context for 20-25 requirements. Asking pays off directly. For tasks where reference agents (built by an engineer) scored 95–100% — the builds that asked zero questions scored 5%, one question 15%, two questions 25%, and so on.

Quelle & Methodik

Diese Standpunkte sind mit ihren Originalquellen verknüpft. Paraphrasen sind gekennzeichnet und keine wörtlichen Zitate.

Transkript oder Quellenmaterial öffnen (wird in einem neuen Tab geöffnet)Ein Problem melden