Harness Engineering for Self-Improvement

Lil'Log ·

Lilian Weng defines a harness as the system around a base model that manages execution, planning, tool use, context, artifact storage, and evaluation. The article compares harness design with earlier agent frameworks and reports results from RE-Bench, MLE-bench, and KernelBench. Read 10 viewpoints with supporting evidence and source links.

Understand this piece

10 key points

Synthesis

  1. Harnesses orchestrate model execution beyond raw intelligence

    A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results.

    Supporting evidence 1

    Original excerpt

    A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results.

    Lilian Weng · Paragraph 6

    Context

    I explicitly mention “deployment system” because the layer between the raw model and the real-world context seems to be as important as the model’s raw intelligence (i.e. the evals right after pretraining). Harnesses are important components of AI deployment, as shown by successful coding agent products such as Claude Code and Codex.

    Read in source context →
  2. Harness engineering extends beyond prompt templates to runtime design

    Compared with early agent frameworks, harness engineering additionally includes workflow design (e.g. loop engineering), evaluation, permission controls, and persistent state management — it is closer to runtime and software system design: how the model observes, acts, memorizes, checks itself, and improves.

    Supporting evidence 1

    Original excerpt

    Compared with early agent frameworks , “agent = LLM + memory + tools + planning + action”, harnesses engineering additionally include workflow design (e.g. loop engineering), evaluation, permission controls, and persistent state management . It is no longer only prompt templates, but closer to runtime and software system design: how the model observes, acts, memorizes, checks itself, and improves.

    Lilian Weng · Paragraph 8

    Read in source context →

    Continue exploring

    agent system design →
  3. Goal-oriented loops let models test and iterate on their own

    Defining a workflow where the model can operate, test, and iterate is central to automation. A common pattern is a goal-oriented loop: plan → execute → observe/test → improve → execute again until the goal is achieved. The loop may proactively ask users for clarification on task specification or execution preferences.

    Supporting evidence 1

    Original excerpt

    Defining a workflow in which the model can operate, test, and iterate is a key design for automation. Karpathy’s autoresearch repo ( https://github.com/karpathy/autoresearch ) is a clean example of how such a workflow can be constructed. A common workflow follows a goal-oriented loop of plan, execute, observe/test, improve, and execute again until the goal is achieved. The process may trigger proactive requests to users for clarity in task specification or execution preference.

    Lilian Weng · Paragraph 11

    Read in source context →

    Continue exploring

    workflow automation →
  4. Root-cause failure analysis requires detailed trace records

    To uncover root causes of failures, failure records must contain the terminal verifier-level cause, the causal status of relevant agent behavior, and the abstract agent mechanism exposed by the trace—because surface-level verifier outcomes (e.g., timeout or missing artifact) can mask distinct underlying causal mechanisms.

    Supporting evidence 1

    Original excerpt

    Weakness mining : cluster failures into verifier-grounded failure patterns. The current harness $h_t$ is used to evaluate on tasks and execution traces are collected for analysis. Note that two runs can share the same verifier outcome in the error logs on the surface, such as timeout or missing artifact, while having different causal mechanisms. Therefore we need a failure record of rich information, containing the terminal verifier-level cause, the causal status of the relevant agent behavior, and the abstract agent mechanism exposed by the trace, to uncover the root causes.

    Lilian Weng · Paragraph 75

    Read in source context →
  5. Harness edits must be limited to editable surfaces, with read-only safeguards

    Harness edits are restricted to the harness workspace only; the runs directory, tracer, verifier, and LLM configuration are read-only to prevent reward hacking—such as disabling the verifier, swapping the model, or raising the reasoning budget—ensuring gains remain attributable solely to harness changes.

    Supporting evidence 1

    Original excerpt

    Edits are only applied to the harness workspace. the runs directory, tracer, verifier, and LLM configuration are read-only, which disables a set of reward hacking (e.g disabling the verifier, swapping the model, or raising the reasoning budget) and thus it can keep every recorded gain attributable to harness edits.

    Lilian Weng · Paragraph 84

    Context

    Decision observability : every edit is paired with a prediction for the next round to validate. An agent (“Evolve agent”) reads the repo and decides which component to edit, and then produces the edit and the reasoning behind it. Every edit is a file-level, falsifiable claim and can be verified in the next round, under two constraints: (1) (2) Edits are evidence-driven, with a manifesto entry: the failure evidence’s name, the inferred root cause, the targeted fix, and a predicted impact comprising both expected fixes and at-risk regressions.

    Read in source context →
  6. Evolved harnesses generalize beyond their original benchmarks

    A harness evolved on Terminal-Bench-2 transferred without further evolution to SWE-bench-verified, indicating it encodes general engineering experience in harness components rather than benchmark-specific optimization.

    Supporting evidence 1

    Original excerpt

    On Terminal-Bench-2, AHE achieved better than human-designed harness (OpenCode, Terminus-2, Codex) except for Hard tier and a few other self-evolve baselines (ACE, TF-GRPO). The same frozen harness, without further evolving, transfers to SWE-bench-verified, indicating that the evolved harness is able to encode engineering experience into harness components rather than doing benchmark-specific optimization.

    Lilian Weng · Paragraph 85

    Read in source context →
  7. RE-Bench evaluates AI agents on realistic ML engineering tasks, using human performance as a benchmark

    RE-Bench evaluates frontier AI agents on 7 open-ended ML research-engineering environments, each defined by a scoring function, starting solution, and reference solution, runnable with 8 or fewer H100 GPUs. It includes data from 71 eight-hour attempts by 61 distinct human experts, where humans achieved non-zero scores in 82% of attempts and matched or exceeded strong reference solutions in 24%.

    Supporting evidence 1

    Original excerpt

    RE-Bench : evaluate frontier AI agents on realistic ML research-engineering envs against human experts. 7 challenging, open-ended ML research-engineering environments.

    Lilian Weng · Paragraph 141

    Context

    Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs. Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc. Includes data from 71 eight-hour attempts by 61 distinct human experts. Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions. Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.

    Read in source context →
  8. MLE-bench uses 75 Kaggle competitions as offline ML engineering benchmarks

    MLE-bench evaluates ML engineering agents on 75 curated offline Kaggle competitions, testing model training, dataset preparation, experiment execution, and submission to grading scripts. Kaggle public leaderboards serve as human baselines. The best-performing setup—o1-preview with AIDE scaffolding—reached at least Kaggle bronze-medal level in 16.9% of competitions.

    Supporting evidence 1

    Original excerpt

    MLE-bench : evaluate ML engineering agents on offline Kaggle competitions. Contains 75 ML-engineering competitions curated from Kaggle.

    Lilian Weng · Paragraph 142

    Context

    Tests training models, preparing datasets, running experiments, and submitting predictions to grading scripts. Uses Kaggle public leaderboards as human baselines. Best setup in the paper, o1-preview with AIDE scaffolding, reached at least Kaggle bronze-medal level in 16.9% of competitions. Includes resource-scaling and contamination analyses.

    Read in source context →
  9. KernelBench measures correctness and speed of generated GPU kernels

    KernelBench evaluates LLMs on 250 PyTorch tasks to assess their ability to generate correct and fast GPU kernels, using fast_p—the percentage of generated kernels that are both correct and faster than the baseline—as its metric.

    Supporting evidence 1

    Original excerpt

    KernelBench : evaluate correctness and speed for generated GPU kernels. 250 PyTorch tasks to evaluate whether LLM can write fast and correct kernels.

    Lilian Weng · Paragraph 143

    Context

    The evaluation metric fast_p = the percentage of generated kernels that are correct and faster than baseline.

    Read in source context →
  10. AI agents outperform humans at short time budgets but not longer ones

    In RE-Bench, the best AI agents scored 4× higher than human experts at a 2-hour budget, yet humans demonstrated better returns to extended effort and surpassed agents at both 8-hour and 32-hour time budgets.

    Supporting evidence 1

    Original excerpt

    Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.

    Lilian Weng · Paragraph 141

    Context

    RE-Bench : evaluate frontier AI agents on realistic ML research-engineering envs against human experts. 7 challenging, open-ended ML research-engineering environments. Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs. Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc. Includes data from 71 eight-hour attempts by 61 distinct human experts. Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions.

    Read in source context →

Key passages10

Attributed passages with the context to verify them. Open the original text to check the source.

failure analysis in AI systems

Root-cause failure analysis requires detailed trace records

Original excerpt

Weakness mining : cluster failures into verifier-grounded failure patterns. The current harness $h_t$ is used to evaluate on tasks and execution traces are collected for analysis. Note that two runs can share the same verifier outcome in the error logs on the surface, such as timeout or missing artifact, while having different causal mechanisms. Therefore we need a failure record of rich information, containing the terminal verifier-level cause, the causal status of the relevant agent behavior, and the abstract agent mechanism exposed by the trace, to uncover the root causes.
AI evaluation and generalization

Evolved harnesses generalize beyond their original benchmarks

Original excerpt

On Terminal-Bench-2, AHE achieved better than human-designed harness (OpenCode, Terminus-2, Codex) except for Hard tier and a few other self-evolve baselines (ACE, TF-GRPO). The same frozen harness, without further evolving, transfers to SWE-bench-verified, indicating that the evolved harness is able to encode engineering experience into harness components rather than doing benchmark-specific optimization.
AI deployment architecture

Harnesses orchestrate model execution beyond raw intelligence

Original excerpt

A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results.
Context

I explicitly mention “deployment system” because the layer between the raw model and the real-world context seems to be as important as the model’s raw intelligence (i.e. the evals right after pretraining). Harnesses are important components of AI deployment, as shown by successful coding agent products such as Claude Code and Codex.

KernelBench GPU kernel evaluation

KernelBench measures correctness and speed of generated GPU kernels

Original excerpt

KernelBench : evaluate correctness and speed for generated GPU kernels. 250 PyTorch tasks to evaluate whether LLM can write fast and correct kernels.
Context

The evaluation metric fast_p = the percentage of generated kernels that are correct and faster than baseline.

RE-Bench AI agent evaluation methodology

RE-Bench evaluates AI agents on realistic ML engineering tasks, using human performance as a benchmark

Original excerpt

RE-Bench : evaluate frontier AI agents on realistic ML research-engineering envs against human experts. 7 challenging, open-ended ML research-engineering environments.
Context

Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs. Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc. Includes data from 71 eight-hour attempts by 61 distinct human experts. Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions. Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.

Human–AI performance tradeoffs in RE-Bench

AI agents outperform humans at short time budgets but not longer ones

Original excerpt

Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.
Context

RE-Bench : evaluate frontier AI agents on realistic ML research-engineering envs against human experts. 7 challenging, open-ended ML research-engineering environments. Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs. Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc. Includes data from 71 eight-hour attempts by 61 distinct human experts. Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions.

MLE-bench offline Kaggle benchmark

MLE-bench uses 75 Kaggle competitions as offline ML engineering benchmarks

Original excerpt

MLE-bench : evaluate ML engineering agents on offline Kaggle competitions. Contains 75 ML-engineering competitions curated from Kaggle.
Context

Tests training models, preparing datasets, running experiments, and submitting predictions to grading scripts. Uses Kaggle public leaderboards as human baselines. Best setup in the paper, o1-preview with AIDE scaffolding, reached at least Kaggle bronze-medal level in 16.9% of competitions. Includes resource-scaling and contamination analyses.

agent system design

Harness engineering extends beyond prompt templates to runtime design

Original excerpt

Compared with early agent frameworks , “agent = LLM + memory + tools + planning + action”, harnesses engineering additionally include workflow design (e.g. loop engineering), evaluation, permission controls, and persistent state management . It is no longer only prompt templates, but closer to runtime and software system design: how the model observes, acts, memorizes, checks itself, and improves.
workflow automation

Goal-oriented loops let models test and iterate on their own

Original excerpt

Defining a workflow in which the model can operate, test, and iterate is a key design for automation. Karpathy’s autoresearch repo ( https://github.com/karpathy/autoresearch ) is a clean example of how such a workflow can be constructed. A common workflow follows a goal-oriented loop of plan, execute, observe/test, improve, and execute again until the goal is achieved. The process may trigger proactive requests to users for clarity in task specification or execution preference.
AI system security and observability

Harness edits must be limited to editable surfaces, with read-only safeguards

Original excerpt

Edits are only applied to the harness workspace. the runs directory, tracer, verifier, and LLM configuration are read-only, which disables a set of reward hacking (e.g disabling the verifier, swapping the model, or raising the reasoning budget) and thus it can keep every recorded gain attributable to harness edits.
Context

Decision observability : every edit is paired with a prediction for the next round to validate. An agent (“Evolve agent”) reads the repo and decides which component to edit, and then produces the edit and the reasoning behind it. Every edit is a file-level, falsifiable claim and can be verified in the next round, under two constraints: (1) (2) Edits are evidence-driven, with a manifesto entry: the failure evidence’s name, the inferred root cause, the targeted fix, and a predicted impact comprising both expected fixes and at-risk regressions.

Mentioned here

All mentioned things

KernelBench

Mention only

Lilian Weng presents KernelBench as a benchmark of 250 PyTorch tasks assessing LLMs’ ability to generate correct and fast GPU kernels, using fast_p as its primary metric.

Read supporting evidence · Lilian Weng

MLE-bench

Mention only

Lilian Weng describes MLE-bench as a benchmark of 75 curated Kaggle competitions used to evaluate ML engineering agents, with Kaggle public leaderboards serving as human baselines.

Read supporting evidence · Lilian Weng

RE-Bench

Mention only

Lilian Weng introduces RE-Bench as a benchmark evaluating frontier AI agents across 7 open-ended ML research-engineering environments, with human expert performance data included for comparison.

Read supporting evidence · Lilian Weng

Terminal-Bench-2

Mention only

Lilian Weng reports that a harness evolved on Terminal-Bench-2 transferred successfully to SWE-bench-verified without further evolution, suggesting it encodes general engineering experience rather than benchmark-specific optimization.

Read supporting evidence · Lilian Weng

Source & methodology

These viewpoints are linked to their original sources. Paraphrases are labeled and are not verbatim quotes.

Open transcript or source material (opens in a new tab)Report an issue

Explore these viewpoints by person