超-𝜏-基准(Hyper-𝜏-bench):评估能构建智能体的智能体

Sierra Blog ·

本文对比了对客户服务智能体的评估与对构建智能体的模型的评估。文中报告了独立运行任务与工程师协同任务的性能,并将客户提问行为与基准任务的成功率关联起来。 阅读 3 条观点,查看支持证据与原始来源。

Ben Shi, Keshav Dhandhania

理解这篇

3 个要点

综合解读

  1. 从客户服务智能体到智能体构建者

    该领域已超越评估模型能否作为可靠的客户服务智能体(𝜏-基准最初的焦点),转而评估模型自身能否构建此类智能体。

    支持这项说法 1

    我们于 2024 年构建了 𝜏-基准,旨在回答当时看似新颖的一个问题:模型能否作为可靠的客户服务智能体?如今这已成为基本门槛。更难的问题是:究竟谁在构建这些智能体?答案正日益变为——模型自身。

    Ben Shi, Keshav Dhandhania · 段落 1

    原始摘录
    We built 𝜏-bench in 2024 to answer a question that felt novel at the time: Can a model act as a reliable customer service agent? That’s table stakes now. The harder question is, who’s building the agent in the first place? Increasingly, it’s the models themselves.
    回到原文语境 →
  2. 报告了在工程师上下文下的成功率

    作者报告称,其最佳独立配置(在 Claude Code 中启用最大推理的 Claude Opus 5)在预留任务上的成功率为 23.9%。当与具备深度上下文的工程师配合时,同类模型在这些任务上的成功率达到 82.2%。

    支持这项说法 1

    独立运行时,我们最佳配置(Claude Opus 5 模型,启用最大推理能力,在 Claude Code 环境中运行)仅在预留评估任务中达成 23.9% 的通过率。当与具备深度上下文知识的工程师协同工作时,同一类模型在相同任务上的通过率达到 82.2%。

    Ben Shi, Keshav Dhandhania · 段落 6

    原始摘录
    Working alone, our best configuration — Claude Opus 5 (max reasoning) running in Claude Code — passes just 23.9% of the held-out evaluation tasks. Paired with an engineer with deep context, the same class of model reaches 82.2% on the same tasks.
    回到原文语境 →
  3. 向客户提问与智能体构建成功率高度相关

    在模拟客户独占 20–25 项需求上下文的任务中,未提问的开发智能体得分为 5%,提一个问题的得分为 15%,提两个问题的得分为 25%,依此类推——这表明客户访谈能带来直接回报。

    支持这项说法 1

    它们并未访谈客户。对于客户独占 20–25 项需求上下文的任务,开发者最多仅提出 4 个问题。提问能直接带来收益:在参考智能体(由工程师构建)得分为 95–100% 的任务中,未提问的构建结果得分为 5%,提问 1 次得分为 15%,提问 2 次得分为 25%,依此类推。

    Ben Shi, Keshav Dhandhania · 段落 9

    原始摘录
    They don’t interview the client. Developers only asked four questions at most for tasks where the client had sole context for 20-25 requirements. Asking pays off directly. For tasks where reference agents (built by an engineer) scored 95–100% — the builds that asked zero questions scored 5%, one question 15%, two questions 25%, and so on.
    回到原文语境 →

关键段落3

带明确归属与语境的原文片段。打开原始文本核查出处。

智能体评估范围

从客户服务智能体到智能体构建者

我们于 2024 年构建了 𝜏-基准,旨在回答当时看似新颖的一个问题:模型能否作为可靠的客户服务智能体?如今这已成为基本门槛。更难的问题是:究竟谁在构建这些智能体?答案正日益变为——模型自身。

原始摘录
We built 𝜏-bench in 2024 to answer a question that felt novel at the time: Can a model act as a reliable customer service agent? That’s table stakes now. The harder question is, who’s building the agent in the first place? Increasingly, it’s the models themselves.
客户访谈行为

向客户提问与智能体构建成功率高度相关

它们并未访谈客户。对于客户独占 20–25 项需求上下文的任务,开发者最多仅提出 4 个问题。提问能直接带来收益:在参考智能体(由工程师构建)得分为 95–100% 的任务中,未提问的构建结果得分为 5%,提问 1 次得分为 15%,提问 2 次得分为 25%,依此类推。

原始摘录
They don’t interview the client. Developers only asked four questions at most for tasks where the client had sole context for 20-25 requirements. Asking pays off directly. For tasks where reference agents (built by an engineer) scored 95–100% — the builds that asked zero questions scored 5%, one question 15%, two questions 25%, and so on.
人机协作性能

报告了在工程师上下文下的成功率

独立运行时,我们最佳配置(Claude Opus 5 模型,启用最大推理能力,在 Claude Code 环境中运行)仅在预留评估任务中达成 23.9% 的通过率。当与具备深度上下文知识的工程师协同工作时,同一类模型在相同任务上的通过率达到 82.2%。

原始摘录
Working alone, our best configuration — Claude Opus 5 (max reasoning) running in Claude Code — passes just 23.9% of the held-out evaluation tasks. Paired with an engineer with deep context, the same class of model reaches 82.2% on the same tasks.

来源与研究方法

这些观点均关联原始来源。转述已明确标注,不作为逐字原话展示。

打开转录或来源材料 (在新标签页中打开)报告问题