From customer service agent to agent builder
The field has moved beyond evaluating whether models can act as reliable customer service agents (𝜏-bench's original focus) to evaluating whether models can themselves build such agents.
Stützende Belege
Hyper-𝜏-bench: Evaluating agents that build agents
Originalauszug
We built 𝜏-bench in 2024 to answer a question that felt novel at the time: Can a model act as a reliable customer service agent? That’s table stakes now. The harder question is, who’s building the agent in the first place? Increasingly, it’s the models themselves.