ÖFFENTLICHE AUSGEDRÜCKTE MEINUNGEN

Lilian Weng

1 Quellen · 10 Standpunkte · 10 Themen

Inhalt aktualisiert:

Lilian Weng zu agent system design, AI deployment architecture, AI evaluation and generalization. Entdecke 10 Standpunkte nach Thema, mit Belegen aus 1 Quelle.

Zusammenhänge erkunden

Standpunkte nach Thema

Zugeordnete Standpunkte nach Veröffentlichungsdatum der Quelle. Eine Momentaufnahme dieser Beiträge, keine abschließende Darstellung persönlicher Überzeugungen.

Übersetzungen dienen dem Leseverständnis; die Originalauszüge bleiben die Quellenevidence.

AI deployment architecture

Thema ansehen

Harnesses orchestrate model execution beyond raw intelligence

A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results.

Stützende Belege

Harness Engineering for Self-Improvement

Originalauszug

A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results.
Kontext

I explicitly mention “deployment system” because the layer between the raw model and the real-world context seems to be as important as the model’s raw intelligence (i.e. the evals right after pretraining). Harnesses are important components of AI deployment, as shown by successful coding agent products such as Claude Code and Codex.

Lilian Weng
Erkenntnisse teilen

agent system design

Thema ansehen

Harness engineering extends beyond prompt templates to runtime design

Compared with early agent frameworks, harness engineering additionally includes workflow design (e.g. loop engineering), evaluation, permission controls, and persistent state management — it is closer to runtime and software system design: how the model observes, acts, memorizes, checks itself, and improves.

Stützende Belege

Harness Engineering for Self-Improvement

Originalauszug

Compared with early agent frameworks , “agent = LLM + memory + tools + planning + action”, harnesses engineering additionally include workflow design (e.g. loop engineering), evaluation, permission controls, and persistent state management . It is no longer only prompt templates, but closer to runtime and software system design: how the model observes, acts, memorizes, checks itself, and improves.
Lilian Weng
Erkenntnisse teilen

workflow automation

Thema ansehen

Goal-oriented loops let models test and iterate on their own

Defining a workflow where the model can operate, test, and iterate is central to automation. A common pattern is a goal-oriented loop: plan → execute → observe/test → improve → execute again until the goal is achieved. The loop may proactively ask users for clarification on task specification or execution preferences.

Stützende Belege

Harness Engineering for Self-Improvement

Originalauszug

Defining a workflow in which the model can operate, test, and iterate is a key design for automation. Karpathy’s autoresearch repo ( https://github.com/karpathy/autoresearch ) is a clean example of how such a workflow can be constructed. A common workflow follows a goal-oriented loop of plan, execute, observe/test, improve, and execute again until the goal is achieved. The process may trigger proactive requests to users for clarity in task specification or execution preference.
Lilian Weng
Erkenntnisse teilen

failure analysis in AI systems

Thema ansehen

Root-cause failure analysis requires detailed trace records

To uncover root causes of failures, failure records must contain the terminal verifier-level cause, the causal status of relevant agent behavior, and the abstract agent mechanism exposed by the trace—because surface-level verifier outcomes (e.g., timeout or missing artifact) can mask distinct underlying causal mechanisms.

Stützende Belege

Harness Engineering for Self-Improvement

Originalauszug

Weakness mining : cluster failures into verifier-grounded failure patterns. The current harness $h_t$ is used to evaluate on tasks and execution traces are collected for analysis. Note that two runs can share the same verifier outcome in the error logs on the surface, such as timeout or missing artifact, while having different causal mechanisms. Therefore we need a failure record of rich information, containing the terminal verifier-level cause, the causal status of the relevant agent behavior, and the abstract agent mechanism exposed by the trace, to uncover the root causes.
Lilian Weng
Erkenntnisse teilen

AI system security and observability

Thema ansehen

Harness edits must be limited to editable surfaces, with read-only safeguards

Harness edits are restricted to the harness workspace only; the runs directory, tracer, verifier, and LLM configuration are read-only to prevent reward hacking—such as disabling the verifier, swapping the model, or raising the reasoning budget—ensuring gains remain attributable solely to harness changes.

Stützende Belege

Harness Engineering for Self-Improvement

Originalauszug

Edits are only applied to the harness workspace. the runs directory, tracer, verifier, and LLM configuration are read-only, which disables a set of reward hacking (e.g disabling the verifier, swapping the model, or raising the reasoning budget) and thus it can keep every recorded gain attributable to harness edits.
Kontext

Decision observability : every edit is paired with a prediction for the next round to validate. An agent (“Evolve agent”) reads the repo and decides which component to edit, and then produces the edit and the reasoning behind it. Every edit is a file-level, falsifiable claim and can be verified in the next round, under two constraints: (1) (2) Edits are evidence-driven, with a manifesto entry: the failure evidence’s name, the inferred root cause, the targeted fix, and a predicted impact comprising both expected fixes and at-risk regressions.

Lilian Weng
Erkenntnisse teilen

AI evaluation and generalization

Thema ansehen

Evolved harnesses generalize beyond their original benchmarks

A harness evolved on Terminal-Bench-2 transferred without further evolution to SWE-bench-verified, indicating it encodes general engineering experience in harness components rather than benchmark-specific optimization.

Stützende Belege

Harness Engineering for Self-Improvement

Originalauszug

On Terminal-Bench-2, AHE achieved better than human-designed harness (OpenCode, Terminus-2, Codex) except for Hard tier and a few other self-evolve baselines (ACE, TF-GRPO). The same frozen harness, without further evolving, transfers to SWE-bench-verified, indicating that the evolved harness is able to encode engineering experience into harness components rather than doing benchmark-specific optimization.
Lilian Weng
Erkenntnisse teilen

RE-Bench AI agent evaluation methodology

Thema ansehen

RE-Bench evaluates AI agents on realistic ML engineering tasks, using human performance as a benchmark

RE-Bench evaluates frontier AI agents on 7 open-ended ML research-engineering environments, each defined by a scoring function, starting solution, and reference solution, runnable with 8 or fewer H100 GPUs. It includes data from 71 eight-hour attempts by 61 distinct human experts, where humans achieved non-zero scores in 82% of attempts and matched or exceeded strong reference solutions in 24%.

Stützende Belege

Harness Engineering for Self-Improvement

Originalauszug

RE-Bench : evaluate frontier AI agents on realistic ML research-engineering envs against human experts. 7 challenging, open-ended ML research-engineering environments.
Kontext

Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs. Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc. Includes data from 71 eight-hour attempts by 61 distinct human experts. Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions. Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.

Lilian Weng
Erkenntnisse teilen

MLE-bench offline Kaggle benchmark

Thema ansehen

MLE-bench uses 75 Kaggle competitions as offline ML engineering benchmarks

MLE-bench evaluates ML engineering agents on 75 curated offline Kaggle competitions, testing model training, dataset preparation, experiment execution, and submission to grading scripts. Kaggle public leaderboards serve as human baselines. The best-performing setup—o1-preview with AIDE scaffolding—reached at least Kaggle bronze-medal level in 16.9% of competitions.

Stützende Belege

Harness Engineering for Self-Improvement

Originalauszug

MLE-bench : evaluate ML engineering agents on offline Kaggle competitions. Contains 75 ML-engineering competitions curated from Kaggle.
Kontext

Tests training models, preparing datasets, running experiments, and submitting predictions to grading scripts. Uses Kaggle public leaderboards as human baselines. Best setup in the paper, o1-preview with AIDE scaffolding, reached at least Kaggle bronze-medal level in 16.9% of competitions. Includes resource-scaling and contamination analyses.

Lilian Weng
Erkenntnisse teilen

KernelBench GPU kernel evaluation

Thema ansehen

KernelBench measures correctness and speed of generated GPU kernels

KernelBench evaluates LLMs on 250 PyTorch tasks to assess their ability to generate correct and fast GPU kernels, using fast_p—the percentage of generated kernels that are both correct and faster than the baseline—as its metric.

Stützende Belege

Harness Engineering for Self-Improvement

Originalauszug

KernelBench : evaluate correctness and speed for generated GPU kernels. 250 PyTorch tasks to evaluate whether LLM can write fast and correct kernels.
Kontext

The evaluation metric fast_p = the percentage of generated kernels that are correct and faster than baseline.

Lilian Weng
Erkenntnisse teilen

Human–AI performance tradeoffs in RE-Bench

Thema ansehen

AI agents outperform humans at short time budgets but not longer ones

In RE-Bench, the best AI agents scored 4× higher than human experts at a 2-hour budget, yet humans demonstrated better returns to extended effort and surpassed agents at both 8-hour and 32-hour time budgets.

Stützende Belege

Harness Engineering for Self-Improvement

Originalauszug

Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.
Kontext

RE-Bench : evaluate frontier AI agents on realistic ML research-engineering envs against human experts. 7 challenging, open-ended ML research-engineering environments. Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs. Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc. Includes data from 71 eight-hour attempts by 61 distinct human experts. Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions.

Lilian Weng
Erkenntnisse teilen

Aussagen nach Quelldatum1

Die Aussagen sind nach dem Veröffentlichungsdatum der Originalquelle geordnet; unterschiedliche Formulierungen belegen keinen Positionswechsel.