Hyper-๐œ-bench: Evaluating agents that build agents

Sierra Blog ยท

The article contrasts evaluating customer-service agents with evaluating models that build agents. It reports standalone and engineer-paired task performance, and relates client questioning to success on benchmark tasks. Read 3 viewpoints with supporting evidence and source links.

Ben Shi, Keshav Dhandhania

Understand this piece

3 key points

Synthesis

  1. From customer service agent to agent builder

    The field has moved beyond evaluating whether models can act as reliable customer service agents (๐œ-bench's original focus) to evaluating whether models can themselves build such agents.

    Supporting evidence 1

    Original excerpt

    We built ๐œ-bench in 2024 to answer a question that felt novel at the time: Can a model act as a reliable customer service agent? Thatโ€™s table stakes now. The harder question is, whoโ€™s building the agent in the first place? Increasingly, itโ€™s the models themselves.

    Ben Shi, Keshav Dhandhania ยท Paragraph 1

    Read in source context โ†’
  2. Reported success with engineer context

    The authors report 23.9% success on held-out tasks for their best standalone configuration, Claude Opus 5 with max reasoning in Claude Code. Paired with an engineer with deep context, the same class of model reaches 82.2% on those tasks.

    Supporting evidence 1

    Original excerpt

    Working alone, our best configuration โ€” Claude Opus 5 (max reasoning) running in Claude Code โ€” passes just 23.9% of the held-out evaluation tasks. Paired with an engineer with deep context, the same class of model reaches 82.2% on the same tasks.

    Ben Shi, Keshav Dhandhania ยท Paragraph 6

    Read in source context โ†’
  3. Client questioning strongly correlates with agent-building success

    On tasks where the simulated client held sole context for 20โ€“25 requirements, developer agents that asked zero questions scored 5%, one question 15%, two questions 25%, and so on โ€” demonstrating a direct payoff from client interviews.

    Supporting evidence 1

    Original excerpt

    They donโ€™t interview the client. Developers only asked four questions at most for tasks where the client had sole context for 20-25 requirements. Asking pays off directly. For tasks where reference agents (built by an engineer) scored 95โ€“100% โ€” the builds that asked zero questions scored 5%, one question 15%, two questions 25%, and so on.

    Ben Shi, Keshav Dhandhania ยท Paragraph 9

    Read in source context โ†’

Key passages3

Attributed passages with the context to verify them. Open the original text to check the source.

agent evaluation scope

From customer service agent to agent builder

Original excerpt

We built ๐œ-bench in 2024 to answer a question that felt novel at the time: Can a model act as a reliable customer service agent? Thatโ€™s table stakes now. The harder question is, whoโ€™s building the agent in the first place? Increasingly, itโ€™s the models themselves.
client interviewing behavior

Client questioning strongly correlates with agent-building success

Original excerpt

They donโ€™t interview the client. Developers only asked four questions at most for tasks where the client had sole context for 20-25 requirements. Asking pays off directly. For tasks where reference agents (built by an engineer) scored 95โ€“100% โ€” the builds that asked zero questions scored 5%, one question 15%, two questions 25%, and so on.
human-AI collaboration performance

Reported success with engineer context

Original excerpt

Working alone, our best configuration โ€” Claude Opus 5 (max reasoning) running in Claude Code โ€” passes just 23.9% of the held-out evaluation tasks. Paired with an engineer with deep context, the same class of model reaches 82.2% on the same tasks.

Source & methodology

These viewpoints are linked to their original sources. Paraphrases are labeled and are not verbatim quotes.

Open transcript or source material (opens in a new tab)Report an issue