Lilian Weng 指出,在 Terminal-Bench-2 上演化的编排系统未经进一步演化即成功迁移到 SWE-bench-verified,表明其编码的是通用工程经验而非特定基准的优化。
Lilian Weng ·
支持这项说法
用于自我改进的 Harness 工程
在 Terminal-Bench-2 上,AHE 的表现优于人类设计的运行框架(OpenCode、Terminus-2、Codex),仅在 Hard 难度层级及少数几个自演化基线(ACE、TF-GRPO)上例外。同一套冻结的运行框架无需进一步演化即可迁移到 SWE-bench-verified,说明所演化的运行框架将工程经验编码到了运行框架组件中,而非针对特定基准进行优化。
原始摘录
On Terminal-Bench-2, AHE achieved better than human-designed harness (OpenCode, Terminus-2, Codex) except for Hard tier and a few other self-evolve baselines (ACE, TF-GRPO). The same frozen harness, without further evolving, transfers to SWE-bench-verified, indicating that the evolved harness is able to encode engineering experience into harness components rather than doing benchmark-specific optimization.