Hyper-šœ-bench: Evaluating agents that build agents

Sierra Blog Ā·

The article contrasts evaluating customer-service agents with evaluating models that build agents. It reports standalone and engineer-paired task performance, and relates client questioning to success on benchmark tasks. Lee 3 puntos de vista con sus evidencias y enlaces a las fuentes.

Ben Shi, Keshav Dhandhania

De un vistazo

  • From customer service agent to agent builder

    The field has moved beyond evaluating whether models can act as reliable customer service agents (šœ-bench's original focus) to evaluating whether models can themselves build such agents.

    Ver el momento de apoyo · PÔrrafo 1
  • Reported success with engineer context

    The authors report 23.9% success on held-out tasks for their best standalone configuration, Claude Opus 5 with max reasoning in Claude Code. Paired with an engineer with deep context, the same class of model reaches 82.2% on those tasks.

    Ver el momento de apoyo · PÔrrafo 6
  • Client questioning strongly correlates with agent-building success

    On tasks where the simulated client held sole context for 20–25 requirements, developer agents that asked zero questions scored 5%, one question 15%, two questions 25%, and so on — demonstrating a direct payoff from client interviews.

    Ver el momento de apoyo · PÔrrafo 9

Pasajes clave3

Pasajes atribuidos con contexto para verificarlos. Abra el texto original para comprobar la fuente.

agent evaluation scope

From customer service agent to agent builder

Extracto original

We built šœ-bench in 2024 to answer a question that felt novel at the time: Can a model act as a reliable customer service agent? That’s table stakes now. The harder question is, who’s building the agent in the first place? Increasingly, it’s the models themselves.
client interviewing behavior

Client questioning strongly correlates with agent-building success

Extracto original

They don’t interview the client. Developers only asked four questions at most for tasks where the client had sole context for 20-25 requirements. Asking pays off directly. For tasks where reference agents (built by an engineer) scored 95–100% — the builds that asked zero questions scored 5%, one question 15%, two questions 25%, and so on.

Fuente y metodologĆ­a

Estas perspectivas enlazan a sus fuentes originales. Las parƔfrasis estƔn identificadas y no son citas textuales.

Abrir transcripción o material de origen (se abre en una pestaña nueva)Reportar un problema