从客户服务智能体到智能体构建者
我们于 2024 年构建了 𝜏-基准,旨在回答当时看似新颖的一个问题:模型能否作为可靠的客户服务智能体?如今这已成为基本门槛。更难的问题是:究竟谁在构建这些智能体?答案正日益变为——模型自身。
原始摘录
We built 𝜏-bench in 2024 to answer a question that felt novel at the time: Can a model act as a reliable customer service agent? That’s table stakes now. The harder question is, who’s building the agent in the first place? Increasingly, it’s the models themselves.