“在 118 次真实应用构建中,Claude Sonnet 5.5 生成的应用评分与 Opus 5 持平;平均每次构建仅需 3.6 轮迭代,而 Opus 5 需要 7.7 轮。在我们比较的所有模型中,它的工具调用失败次数最少;也很少在构建中途暂停并询问用户,因此因等待人工响应而停滞的构建更少。”
原始摘录
“Across 118 real app builds, Claude Sonnet 5.5 produced apps that scored level with Opus 5. It got there in 3.6 iterations per build on average, where Opus 5 took 7.7. It had the fewest failed tool calls of any model we compared. It also rarely stopped mid-build to ask the user a question, so fewer builds stall waiting on someone to answer.”
“In Epic’s early testing, Claude Sonnet 5.5 cleared the same quality bar you’d expect from a higher-tier model, holding up on a system design audit and a data flow review. The new model managed tens of thousands of lines of code for gameplay system architecture, kept responses snappy, handled multi-hour tasks, and delivered with less prescriptive prompting.”
“Without changing any of our prompts, Claude Sonnet 5.5 did better than Sonnet 5 on almost all of our offline Slackbot evals, in fewer steps and with about 14% fewer output tokens. When someone gives Slackbot a task, quality and speed are what matter most, and Sonnet 5.5 allows Slackbot to deliver better outcomes for users, faster.”
“我们向 Claude Sonnet 5.5 输入了数百个真实的客服场景(涵盖回复与升级请求)。相比我们当前产线上所用的 Claude 模型,它所做错误决策更少,工单解决速度更快。工单处理速度提升 20%,使我们的客户无需等待即可获得所需帮助。”
原始摘录
“We fed Claude Sonnet 5.5 hundreds of real support use cases across replies and escalation requests. It made fewer wrong decisions and resolved tickets faster than the Claude models we use in production today. Tickets were processed 20% faster, getting our customers the help they need without the wait.”