UN TEMA, EN CONTEXTO

AI evaluation and generalization

Judgments in this source concerning AI evaluation and generalization. Explora 1 punto de vista con evidencias de 1 fuente.

1 personas · 1 fuentes · 1 opiniones expresadas

Contenido actualizado:

Explorar conexiones ↗

Mapa de perspectivas

Explore por persona. Seleccione dos o tres para compararlas.

1 personas · 1 fuentes · 1 opiniones expresadas

Lilian Weng

Evolved harnesses generalize beyond their original benchmarks

A harness evolved on Terminal-Bench-2 transferred without further evolution to SWE-bench-verified, indicating it encodes general engineering experience in harness components rather than benchmark-specific optimization.

Evidencia a favor

Harness Engineering for Self-Improvement

Extracto original

On Terminal-Bench-2, AHE achieved better than human-designed harness (OpenCode, Terminus-2, Codex) except for Hard tier and a few other self-evolve baselines (ACE, TF-GRPO). The same frozen harness, without further evolving, transfers to SWE-bench-verified, indicating that the evolved harness is able to encode engineering experience into harness components rather than doing benchmark-specific optimization.
Compartir informaciónVerificar esta afirmación

Estas son perspectivas individuales, no una medida de consenso. El material fuente permanece en su idioma original.