UN THÈME, DANS SON CONTEXTE

AI evaluation and generalization

Judgments in this source concerning AI evaluation and generalization. Explorez 1 point de vue avec des éléments tirés de 1 source.

1 personnes · 1 sources · 1 opinions exprimées

Contenu mis à jour:

Explorer les liens ↗

Carte des points de vue

Explorer par personne. Sélectionnez deux ou trois personnes pour les comparer.

1 personnes · 1 sources · 1 opinions exprimées

Lilian Weng

Evolved harnesses generalize beyond their original benchmarks

A harness evolved on Terminal-Bench-2 transferred without further evolution to SWE-bench-verified, indicating it encodes general engineering experience in harness components rather than benchmark-specific optimization.

Éléments favorables

Harness Engineering for Self-Improvement

Extrait original

On Terminal-Bench-2, AHE achieved better than human-designed harness (OpenCode, Terminus-2, Codex) except for Hard tier and a few other self-evolve baselines (ACE, TF-GRPO). The same frozen harness, without further evolving, transfers to SWE-bench-verified, indicating that the evolved harness is able to encode engineering experience into harness components rather than doing benchmark-specific optimization.
Partager un aperçuVérifier cette affirmation

Il s’agit de points de vue individuels, non d’une mesure du consensus. Le matériel source reste dans sa langue d’origine.