Client questioning strongly correlates with agent-building success
On tasks where the simulated client held sole context for 20–25 requirements, developer agents that asked zero questions scored 5%, one question 15%, two questions 25%, and so on — demonstrating a direct payoff from client interviews.
Supporting evidence
Hyper-𝜏-bench: Evaluating agents that build agents
Original excerpt
They don’t interview the client. Developers only asked four questions at most for tasks where the client had sole context for 20-25 requirements. Asking pays off directly. For tasks where reference agents (built by an engineer) scored 95–100% — the builds that asked zero questions scored 5%, one question 15%, two questions 25%, and so on.