Harness Engineering for Self-Improvement

Lil'Log ·

Lilian Weng defines a harness as the system around a base model that manages execution, planning, tool use, context, artifact storage, and evaluation. The article compares harness design with earlier agent frameworks and reports results from RE-Bench, MLE-bench, and KernelBench. Lisez 10 points de vue avec leurs éléments à l’appui et les liens vers les sources.

En un coup d’œil

  • Harnesses orchestrate model execution beyond raw intelligence

    A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results.

    Lire le moment probant · Paragraphe 6
  • Harness engineering extends beyond prompt templates to runtime design

    Compared with early agent frameworks, harness engineering additionally includes workflow design (e.g. loop engineering), evaluation, permission controls, and persistent state management — it is closer to runtime and software system design: how the model observes, acts, memorizes, checks itself, and improves.

    Lire le moment probant · Paragraphe 8
  • Goal-oriented loops let models test and iterate on their own

    Defining a workflow where the model can operate, test, and iterate is central to automation. A common pattern is a goal-oriented loop: plan → execute → observe/test → improve → execute again until the goal is achieved. The loop may proactively ask users for clarification on task specification or execution preferences.

    Lire le moment probant · Paragraphe 11
  • Root-cause failure analysis requires detailed trace records

    To uncover root causes of failures, failure records must contain the terminal verifier-level cause, the causal status of relevant agent behavior, and the abstract agent mechanism exposed by the trace—because surface-level verifier outcomes (e.g., timeout or missing artifact) can mask distinct underlying causal mechanisms.

    Lire le moment probant · Paragraphe 75
  • Harness edits must be limited to editable surfaces, with read-only safeguards

    Harness edits are restricted to the harness workspace only; the runs directory, tracer, verifier, and LLM configuration are read-only to prevent reward hacking—such as disabling the verifier, swapping the model, or raising the reasoning budget—ensuring gains remain attributable solely to harness changes.

    Lire le moment probant · Paragraphe 84
  • Evolved harnesses generalize beyond their original benchmarks

    A harness evolved on Terminal-Bench-2 transferred without further evolution to SWE-bench-verified, indicating it encodes general engineering experience in harness components rather than benchmark-specific optimization.

    Lire le moment probant · Paragraphe 85
  • RE-Bench evaluates AI agents on realistic ML engineering tasks, using human performance as a benchmark

    RE-Bench evaluates frontier AI agents on 7 open-ended ML research-engineering environments, each defined by a scoring function, starting solution, and reference solution, runnable with 8 or fewer H100 GPUs. It includes data from 71 eight-hour attempts by 61 distinct human experts, where humans achieved non-zero scores in 82% of attempts and matched or exceeded strong reference solutions in 24%.

    Lire le moment probant · Paragraphe 141
  • MLE-bench uses 75 Kaggle competitions as offline ML engineering benchmarks

    MLE-bench evaluates ML engineering agents on 75 curated offline Kaggle competitions, testing model training, dataset preparation, experiment execution, and submission to grading scripts. Kaggle public leaderboards serve as human baselines. The best-performing setup—o1-preview with AIDE scaffolding—reached at least Kaggle bronze-medal level in 16.9% of competitions.

    Lire le moment probant · Paragraphe 142
  • KernelBench measures correctness and speed of generated GPU kernels

    KernelBench evaluates LLMs on 250 PyTorch tasks to assess their ability to generate correct and fast GPU kernels, using fast_p—the percentage of generated kernels that are both correct and faster than the baseline—as its metric.

    Lire le moment probant · Paragraphe 143
  • AI agents outperform humans at short time budgets but not longer ones

    In RE-Bench, the best AI agents scored 4× higher than human experts at a 2-hour budget, yet humans demonstrated better returns to extended effort and surpassed agents at both 8-hour and 32-hour time budgets.

    Lire le moment probant · Paragraphe 141

Passages clés10

Passages attribués et accompagnés du contexte nécessaire à leur vérification. Ouvrez le texte original pour vérifier la source.

failure analysis in AI systems

Root-cause failure analysis requires detailed trace records

Extrait original

Weakness mining : cluster failures into verifier-grounded failure patterns. The current harness $h_t$ is used to evaluate on tasks and execution traces are collected for analysis. Note that two runs can share the same verifier outcome in the error logs on the surface, such as timeout or missing artifact, while having different causal mechanisms. Therefore we need a failure record of rich information, containing the terminal verifier-level cause, the causal status of the relevant agent behavior, and the abstract agent mechanism exposed by the trace, to uncover the root causes.
AI evaluation and generalization

Evolved harnesses generalize beyond their original benchmarks

Extrait original

On Terminal-Bench-2, AHE achieved better than human-designed harness (OpenCode, Terminus-2, Codex) except for Hard tier and a few other self-evolve baselines (ACE, TF-GRPO). The same frozen harness, without further evolving, transfers to SWE-bench-verified, indicating that the evolved harness is able to encode engineering experience into harness components rather than doing benchmark-specific optimization.
AI deployment architecture

Harnesses orchestrate model execution beyond raw intelligence

Extrait original

A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results.
Contexte

I explicitly mention “deployment system” because the layer between the raw model and the real-world context seems to be as important as the model’s raw intelligence (i.e. the evals right after pretraining). Harnesses are important components of AI deployment, as shown by successful coding agent products such as Claude Code and Codex.

KernelBench GPU kernel evaluation

KernelBench measures correctness and speed of generated GPU kernels

Extrait original

KernelBench : evaluate correctness and speed for generated GPU kernels. 250 PyTorch tasks to evaluate whether LLM can write fast and correct kernels.
Contexte

The evaluation metric fast_p = the percentage of generated kernels that are correct and faster than baseline.

RE-Bench AI agent evaluation methodology

RE-Bench evaluates AI agents on realistic ML engineering tasks, using human performance as a benchmark

Extrait original

RE-Bench : evaluate frontier AI agents on realistic ML research-engineering envs against human experts. 7 challenging, open-ended ML research-engineering environments.
Contexte

Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs. Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc. Includes data from 71 eight-hour attempts by 61 distinct human experts. Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions. Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.

Human–AI performance tradeoffs in RE-Bench

AI agents outperform humans at short time budgets but not longer ones

Extrait original

Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.
Contexte

RE-Bench : evaluate frontier AI agents on realistic ML research-engineering envs against human experts. 7 challenging, open-ended ML research-engineering environments. Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs. Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc. Includes data from 71 eight-hour attempts by 61 distinct human experts. Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions.

MLE-bench offline Kaggle benchmark

MLE-bench uses 75 Kaggle competitions as offline ML engineering benchmarks

Extrait original

MLE-bench : evaluate ML engineering agents on offline Kaggle competitions. Contains 75 ML-engineering competitions curated from Kaggle.
Contexte

Tests training models, preparing datasets, running experiments, and submitting predictions to grading scripts. Uses Kaggle public leaderboards as human baselines. Best setup in the paper, o1-preview with AIDE scaffolding, reached at least Kaggle bronze-medal level in 16.9% of competitions. Includes resource-scaling and contamination analyses.

agent system design

Harness engineering extends beyond prompt templates to runtime design

Extrait original

Compared with early agent frameworks , “agent = LLM + memory + tools + planning + action”, harnesses engineering additionally include workflow design (e.g. loop engineering), evaluation, permission controls, and persistent state management . It is no longer only prompt templates, but closer to runtime and software system design: how the model observes, acts, memorizes, checks itself, and improves.
workflow automation

Goal-oriented loops let models test and iterate on their own

Extrait original

Defining a workflow in which the model can operate, test, and iterate is a key design for automation. Karpathy’s autoresearch repo ( https://github.com/karpathy/autoresearch ) is a clean example of how such a workflow can be constructed. A common workflow follows a goal-oriented loop of plan, execute, observe/test, improve, and execute again until the goal is achieved. The process may trigger proactive requests to users for clarity in task specification or execution preference.
AI system security and observability

Harness edits must be limited to editable surfaces, with read-only safeguards

Extrait original

Edits are only applied to the harness workspace. the runs directory, tracer, verifier, and LLM configuration are read-only, which disables a set of reward hacking (e.g disabling the verifier, swapping the model, or raising the reasoning budget) and thus it can keep every recorded gain attributable to harness edits.
Contexte

Decision observability : every edit is paired with a prediction for the next round to validate. An agent (“Evolve agent”) reads the repo and decides which component to edit, and then produces the edit and the reasoning behind it. Every edit is a file-level, falsifiable claim and can be verified in the next round, under two constraints: (1) (2) Edits are evidence-driven, with a manifesto entry: the failure evidence’s name, the inferred root cause, the targeted fix, and a predicted impact comprising both expected fixes and at-risk regressions.

Source et méthodologie

Ces points de vue renvoient à leurs sources originales. Les reformulations sont signalées et ne sont pas des citations mot à mot.

Ouvrir la transcription ou les documents sources (s’ouvre dans un nouvel onglet)Signaler un problème