Base44 报告称,在 118 个应用构建任务中所需迭代次数更少
在 Anthropic 发布的一则客户证言中,Base44 公司的加布里埃尔·格林伯格(Gabriel Grinberg)表示,Claude Sonnet 5.5 在 118 个应用构建任务中与 Opus 5 的评分持平,平均每个构建仅需 3.6 轮迭代,而 Opus 5 则需 7.7 轮。
支持这项说法
Anthropic 关于 Claude Sonnet 5.5 的客户证言
“在 118 次真实应用构建中,Claude Sonnet 5.5 生成的应用评分与 Opus 5 持平;平均每次构建仅需 3.6 轮迭代,而 Opus 5 需要 7.7 轮。在我们比较的所有模型中,它的工具调用失败次数最少;也很少在构建中途暂停并询问用户,因此因等待人工响应而停滞的构建更少。”
原始摘录
“Across 118 real app builds, Claude Sonnet 5.5 produced apps that scored level with Opus 5. It got there in 3.6 iterations per build on average, where Opus 5 took 7.7. It had the fewest failed tool calls of any model we compared. It also rarely stopped mid-build to ask the user a question, so fewer builds stall waiting on someone to answer.”