用于自我改进的 Harness 工程
缺陷挖掘(Weakness mining):将失败案例按验证器确认的失败模式聚类。当前运行框架 $h_t$ 用于任务评测,并收集执行轨迹以供分析。需注意:两次运行在错误日志表面可能呈现相同的验证器结果(如超时或缺失产物),但根本成因可能不同。因此,我们需要信息丰富的失败记录,其中须包含终端验证器层面的根本原因、相关智能体行为的因果状态,以及执行轨迹所揭示的抽象智能体机制,从而定位深层根因。
原始摘录
Weakness mining : cluster failures into verifier-grounded failure patterns. The current harness $h_t$ is used to evaluate on tasks and execution traces are collected for analysis. Note that two runs can share the same verifier outcome in the error logs on the surface, such as timeout or missing artifact, while having different causal mechanisms. Therefore we need a failure record of rich information, containing the terminal verifier-level cause, the causal status of the relevant agent behavior, and the abstract agent mechanism exposed by the trace, to uncover the root causes.