A TOPIC, IN CONTEXT

AI evaluation and generalization

Judgments in this source concerning AI evaluation and generalization. Explore 1 viewpoint with evidence from 1 source.

1 people · 1 sources · 1 viewpoints

Content updated:

Explore connections ↗

Viewpoint map

Explore by person. Select two or three to compare.

1 people · 1 sources · 1 viewpoints

Lilian Weng

Evolved harnesses generalize beyond their original benchmarks

A harness evolved on Terminal-Bench-2 transferred without further evolution to SWE-bench-verified, indicating it encodes general engineering experience in harness components rather than benchmark-specific optimization.

Supporting evidence

Harness Engineering for Self-Improvement

Original excerpt

On Terminal-Bench-2, AHE achieved better than human-designed harness (OpenCode, Terminus-2, Codex) except for Hard tier and a few other self-evolve baselines (ACE, TF-GRPO). The same frozen harness, without further evolving, transfers to SWE-bench-verified, indicating that the evolved harness is able to encode engineering experience into harness components rather than doing benchmark-specific optimization.
Share insightCheck this claim

These are individual perspectives, not a measure of consensus. Source material stays in its original language.