Anthropic 关于 Claude Sonnet 5.5 的客户证言

Anthropic News ·

Anthropic 关于 Claude Sonnet 5.5 的发布声明中,包含了来自 Epic Games、Base44、Slack 和 Zendesk 的具名客户证言。这些客户描述了其自身在工程实践、应用构建、助手测试及客服支持等方面的测试情况。相关证言由 Anthropic 精选并发布;其所述结果未经 nafyi 独立验证。 阅读 4 条观点,查看支持证据与原始来源。

理解这篇

4 个要点

综合解读

  1. Epic 报告称,该模型在工程任务中展现出更高层级模型的质量水平

    在 Anthropic 发布的一则客户证言中,Epic Games 的 Daniel Vogel 表示,Sonnet 5.5 在早期系统设计与数据流评审中达到了他所预期的更高层级模型的质量标准,同时在处理长周期工程任务时所需提示词的指导性更弱。

    支持这项说法 1

    “在 Epic 公司的早期测试中,Claude Sonnet 5.5 达到了人们对其高阶模型所预期的质量标准,在系统设计审计和数据流审查中均表现稳健。该新模型可管理数万行游戏系统架构代码,响应迅捷,能处理持续数小时的复杂任务,且所需指令性提示更少。”

    Daniel Vogel · Source text block 30 · customer quotation

    原始摘录
    “In Epic’s early testing, Claude Sonnet 5.5 cleared the same quality bar you’d expect from a higher-tier model, holding up on a system design audit and a data flow review. The new model managed tens of thousands of lines of code for gameplay system architecture, kept responses snappy, handled multi-hour tasks, and delivered with less prescriptive prompting.”
    回到原文语境 →
  2. Base44 报告称,在 118 个应用构建任务中所需迭代次数更少

    在 Anthropic 发布的一则客户证言中,Base44 公司的加布里埃尔·格林伯格(Gabriel Grinberg)表示,Claude Sonnet 5.5 在 118 个应用构建任务中与 Opus 5 的评分持平,平均每个构建仅需 3.6 轮迭代,而 Opus 5 则需 7.7 轮。

    支持这项说法 1

    “在 118 次真实应用构建中,Claude Sonnet 5.5 生成的应用评分与 Opus 5 持平;平均每次构建仅需 3.6 轮迭代,而 Opus 5 需要 7.7 轮。在我们比较的所有模型中,它的工具调用失败次数最少;也很少在构建中途暂停并询问用户,因此因等待人工响应而停滞的构建更少。”

    Gabriel Grinberg · Source text block 34 · customer quotation

    原始摘录
    “Across 118 real app builds, Claude Sonnet 5.5 produced apps that scored level with Opus 5. It got there in 3.6 iterations per build on average, where Opus 5 took 7.7. It had the fewest failed tool calls of any model we compared. It also rarely stopped mid-build to ask the user a question, so fewer builds stall waiting on someone to answer.”
    回到原文语境 →
  3. Slack 报告称,在离线评估中结果更优且输出 token 更少

    在 Anthropic 发布的一则客户证言中,Slack 公司的柯蒂斯·艾伦(Curtis Allen)表示,在几乎全部离线 Slackbot 评估任务中(未更改任何提示词),Sonnet 5.5 表现优于 Sonnet 5,所用步骤更少,输出 token 数量减少约 14%。

    支持这项说法 1

    “在不更改任何提示词的前提下,Claude Sonnet 5.5 在几乎全部离线 Slackbot 评估任务中的表现均优于 Sonnet 5,所用步骤更少,输出 token 数量减少约 14%。当用户向 Slackbot 下达任务时,质量与速度最为关键,而 Sonnet 5.5 使 Slackbot 能够更快地为用户提供更优的结果。”

    Curtis Allen · Source text block 41 · customer quotation

    原始摘录
    “Without changing any of our prompts, Claude Sonnet 5.5 did better than Sonnet 5 on almost all of our offline Slackbot evals, in fewer steps and with about 14% fewer output tokens. When someone gives Slackbot a task, quality and speed are what matter most, and Sonnet 5.5 allows Slackbot to deliver better outcomes for users, faster.”
    回到原文语境 →
  4. Zendesk 表示,错误决策更少,处理速度更快

    在 Anthropic 发布的一则客户证言中,Zendesk 的 Abhinay Kathuria 表示,与 Zendesk 当时使用的 Claude 模型相比,Sonnet 5.5 在数百种支持服务用例中错误决策更少,处理工单速度快了 20%。

    支持这项说法 1

    “我们向 Claude Sonnet 5.5 输入了数百个真实的客服场景(涵盖回复与升级请求)。相比我们当前产线上所用的 Claude 模型,它所做错误决策更少,工单解决速度更快。工单处理速度提升 20%,使我们的客户无需等待即可获得所需帮助。”

    Abhinay Kathuria · Source text block 42 · customer quotation

    原始摘录
    “We fed Claude Sonnet 5.5 hundreds of real support use cases across replies and escalation requests. It made fewer wrong decisions and resolved tickets faster than the Claude models we use in production today. Tickets were processed 20% faster, getting our customers the help they need without the wait.”
    回到原文语境 →

关键段落4

带明确归属与语境的原文片段。打开原始文本核查出处。

应用构建迭代

Base44 报告称,在 118 个应用构建任务中所需迭代次数更少

“在 118 次真实应用构建中,Claude Sonnet 5.5 生成的应用评分与 Opus 5 持平;平均每次构建仅需 3.6 轮迭代,而 Opus 5 需要 7.7 轮。在我们比较的所有模型中,它的工具调用失败次数最少;也很少在构建中途暂停并询问用户,因此因等待人工响应而停滞的构建更少。”

原始摘录
“Across 118 real app builds, Claude Sonnet 5.5 produced apps that scored level with Opus 5. It got there in 3.6 iterations per build on average, where Opus 5 took 7.7. It had the fewest failed tool calls of any model we compared. It also rarely stopped mid-build to ask the user a question, so fewer builds stall waiting on someone to answer.”
助手评估效率

Slack 报告称,在离线评估中结果更优且输出 token 更少

“在不更改任何提示词的前提下,Claude Sonnet 5.5 在几乎全部离线 Slackbot 评估任务中的表现均优于 Sonnet 5,所用步骤更少,输出 token 数量减少约 14%。当用户向 Slackbot 下达任务时,质量与速度最为关键,而 Sonnet 5.5 使 Slackbot 能够更快地为用户提供更优的结果。”

原始摘录
“Without changing any of our prompts, Claude Sonnet 5.5 did better than Sonnet 5 on almost all of our offline Slackbot evals, in fewer steps and with about 14% fewer output tokens. When someone gives Slackbot a task, quality and speed are what matter most, and Sonnet 5.5 allows Slackbot to deliver better outcomes for users, faster.”
客户服务工作流

Zendesk 表示,错误决策更少,处理速度更快

“我们向 Claude Sonnet 5.5 输入了数百个真实的客服场景(涵盖回复与升级请求)。相比我们当前产线上所用的 Claude 模型,它所做错误决策更少,工单解决速度更快。工单处理速度提升 20%,使我们的客户无需等待即可获得所需帮助。”

原始摘录
“We fed Claude Sonnet 5.5 hundreds of real support use cases across replies and escalation requests. It made fewer wrong decisions and resolved tickets faster than the Claude models we use in production today. Tickets were processed 20% faster, getting our customers the help they need without the wait.”
编码任务质量

Epic 报告称,该模型在工程任务中展现出更高层级模型的质量水平

“在 Epic 公司的早期测试中,Claude Sonnet 5.5 达到了人们对其高阶模型所预期的质量标准,在系统设计审计和数据流审查中均表现稳健。该新模型可管理数万行游戏系统架构代码,响应迅捷,能处理持续数小时的复杂任务,且所需指令性提示更少。”

原始摘录
“In Epic’s early testing, Claude Sonnet 5.5 cleared the same quality bar you’d expect from a higher-tier model, holding up on a system design audit and a data flow review. The new model managed tens of thousands of lines of code for gameplay system architecture, kept responses snappy, handled multi-hour tasks, and delivered with less prescriptive prompting.”

这里提到的

全部提及对象

Claude Sonnet 5.5

支持

Base44 的 Gabriel Grinberg 表示,Claude Sonnet 5.5 在 118 次真实应用构建中与 Opus 5 达到同等评分,平均仅需 3.6 次迭代(Opus 5 为 7.7 次),且工具调用失败更少、中途停顿提问也更少。

查看支持证据 · Gabriel Grinberg
支持

Epic Games 的 Daniel Vogel 表示,Claude Sonnet 5.5 在早期系统设计和数据流审查中达到了对更高级别模型的预期质量标准,同时能以更少的指令性提示处理长时间的工程任务。

查看支持证据 · Daniel Vogel
支持

Slack 的 Curtis Allen 表示,Claude Sonnet 5.5 在不更改任何提示词的前提下,在大多数离线 Slackbot 评估中优于 Sonnet 5,步骤更少且输出 token 减少约 14%。

查看支持证据 · Curtis Allen
支持

Zendesk 的 Abhinay Kathuria 表示,Claude Sonnet 5.5 在数百个真实客服场景中错误决策更少,且工单处理速度比 Zendesk 当前部署的 Claude 模型快 20%。

查看支持证据 · Abhinay Kathuria

来源与研究方法

这些观点均关联原始来源。转述已明确标注,不作为逐字原话展示。

打开转录或来源材料 (在新标签页中打开)报告问题

继续了解这些人物的观点