OPINIONS EXPRIMÉES EN PUBLIC

Lilian Weng

1 sources · 10 points de vue · 10 sujets

Contenu mis à jour:

Lilian Weng sur agent system design, AI deployment architecture, AI evaluation and generalization. Explorez 10 points de vue par thème, avec des éléments tirés de 1 source.

Explorer les liens

Points de vue par sujet

Points de vue attribués, classés par date de publication de la source. Un aperçu de ces échanges, sans prétendre définir toutes les convictions de la personne.

Les traductions sont destinées à la lecture ; les extraits originaux restent la source evidence.

AI deployment architecture

Voir ce sujet

Harnesses orchestrate model execution beyond raw intelligence

A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results.

Éléments favorables

Harness Engineering for Self-Improvement

Extrait original

A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results.
Contexte

I explicitly mention “deployment system” because the layer between the raw model and the real-world context seems to be as important as the model’s raw intelligence (i.e. the evals right after pretraining). Harnesses are important components of AI deployment, as shown by successful coding agent products such as Claude Code and Codex.

Lilian Weng
Partager un aperçu

agent system design

Voir ce sujet

Harness engineering extends beyond prompt templates to runtime design

Compared with early agent frameworks, harness engineering additionally includes workflow design (e.g. loop engineering), evaluation, permission controls, and persistent state management — it is closer to runtime and software system design: how the model observes, acts, memorizes, checks itself, and improves.

Éléments favorables

Harness Engineering for Self-Improvement

Extrait original

Compared with early agent frameworks , “agent = LLM + memory + tools + planning + action”, harnesses engineering additionally include workflow design (e.g. loop engineering), evaluation, permission controls, and persistent state management . It is no longer only prompt templates, but closer to runtime and software system design: how the model observes, acts, memorizes, checks itself, and improves.
Lilian Weng
Partager un aperçu

workflow automation

Voir ce sujet

Goal-oriented loops let models test and iterate on their own

Defining a workflow where the model can operate, test, and iterate is central to automation. A common pattern is a goal-oriented loop: plan → execute → observe/test → improve → execute again until the goal is achieved. The loop may proactively ask users for clarification on task specification or execution preferences.

Éléments favorables

Harness Engineering for Self-Improvement

Extrait original

Defining a workflow in which the model can operate, test, and iterate is a key design for automation. Karpathy’s autoresearch repo ( https://github.com/karpathy/autoresearch ) is a clean example of how such a workflow can be constructed. A common workflow follows a goal-oriented loop of plan, execute, observe/test, improve, and execute again until the goal is achieved. The process may trigger proactive requests to users for clarity in task specification or execution preference.
Lilian Weng
Partager un aperçu

failure analysis in AI systems

Voir ce sujet

Root-cause failure analysis requires detailed trace records

To uncover root causes of failures, failure records must contain the terminal verifier-level cause, the causal status of relevant agent behavior, and the abstract agent mechanism exposed by the trace—because surface-level verifier outcomes (e.g., timeout or missing artifact) can mask distinct underlying causal mechanisms.

Éléments favorables

Harness Engineering for Self-Improvement

Extrait original

Weakness mining : cluster failures into verifier-grounded failure patterns. The current harness $h_t$ is used to evaluate on tasks and execution traces are collected for analysis. Note that two runs can share the same verifier outcome in the error logs on the surface, such as timeout or missing artifact, while having different causal mechanisms. Therefore we need a failure record of rich information, containing the terminal verifier-level cause, the causal status of the relevant agent behavior, and the abstract agent mechanism exposed by the trace, to uncover the root causes.
Lilian Weng
Partager un aperçu

AI system security and observability

Voir ce sujet

Harness edits must be limited to editable surfaces, with read-only safeguards

Harness edits are restricted to the harness workspace only; the runs directory, tracer, verifier, and LLM configuration are read-only to prevent reward hacking—such as disabling the verifier, swapping the model, or raising the reasoning budget—ensuring gains remain attributable solely to harness changes.

Éléments favorables

Harness Engineering for Self-Improvement

Extrait original

Edits are only applied to the harness workspace. the runs directory, tracer, verifier, and LLM configuration are read-only, which disables a set of reward hacking (e.g disabling the verifier, swapping the model, or raising the reasoning budget) and thus it can keep every recorded gain attributable to harness edits.
Contexte

Decision observability : every edit is paired with a prediction for the next round to validate. An agent (“Evolve agent”) reads the repo and decides which component to edit, and then produces the edit and the reasoning behind it. Every edit is a file-level, falsifiable claim and can be verified in the next round, under two constraints: (1) (2) Edits are evidence-driven, with a manifesto entry: the failure evidence’s name, the inferred root cause, the targeted fix, and a predicted impact comprising both expected fixes and at-risk regressions.

Lilian Weng
Partager un aperçu

AI evaluation and generalization

Voir ce sujet

Evolved harnesses generalize beyond their original benchmarks

A harness evolved on Terminal-Bench-2 transferred without further evolution to SWE-bench-verified, indicating it encodes general engineering experience in harness components rather than benchmark-specific optimization.

Éléments favorables

Harness Engineering for Self-Improvement

Extrait original

On Terminal-Bench-2, AHE achieved better than human-designed harness (OpenCode, Terminus-2, Codex) except for Hard tier and a few other self-evolve baselines (ACE, TF-GRPO). The same frozen harness, without further evolving, transfers to SWE-bench-verified, indicating that the evolved harness is able to encode engineering experience into harness components rather than doing benchmark-specific optimization.
Lilian Weng
Partager un aperçu

RE-Bench AI agent evaluation methodology

Voir ce sujet

RE-Bench evaluates AI agents on realistic ML engineering tasks, using human performance as a benchmark

RE-Bench evaluates frontier AI agents on 7 open-ended ML research-engineering environments, each defined by a scoring function, starting solution, and reference solution, runnable with 8 or fewer H100 GPUs. It includes data from 71 eight-hour attempts by 61 distinct human experts, where humans achieved non-zero scores in 82% of attempts and matched or exceeded strong reference solutions in 24%.

Éléments favorables

Harness Engineering for Self-Improvement

Extrait original

RE-Bench : evaluate frontier AI agents on realistic ML research-engineering envs against human experts. 7 challenging, open-ended ML research-engineering environments.
Contexte

Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs. Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc. Includes data from 71 eight-hour attempts by 61 distinct human experts. Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions. Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.

Lilian Weng
Partager un aperçu

MLE-bench offline Kaggle benchmark

Voir ce sujet

MLE-bench uses 75 Kaggle competitions as offline ML engineering benchmarks

MLE-bench evaluates ML engineering agents on 75 curated offline Kaggle competitions, testing model training, dataset preparation, experiment execution, and submission to grading scripts. Kaggle public leaderboards serve as human baselines. The best-performing setup—o1-preview with AIDE scaffolding—reached at least Kaggle bronze-medal level in 16.9% of competitions.

Éléments favorables

Harness Engineering for Self-Improvement

Extrait original

MLE-bench : evaluate ML engineering agents on offline Kaggle competitions. Contains 75 ML-engineering competitions curated from Kaggle.
Contexte

Tests training models, preparing datasets, running experiments, and submitting predictions to grading scripts. Uses Kaggle public leaderboards as human baselines. Best setup in the paper, o1-preview with AIDE scaffolding, reached at least Kaggle bronze-medal level in 16.9% of competitions. Includes resource-scaling and contamination analyses.

Lilian Weng
Partager un aperçu

KernelBench GPU kernel evaluation

Voir ce sujet

KernelBench measures correctness and speed of generated GPU kernels

KernelBench evaluates LLMs on 250 PyTorch tasks to assess their ability to generate correct and fast GPU kernels, using fast_p—the percentage of generated kernels that are both correct and faster than the baseline—as its metric.

Éléments favorables

Harness Engineering for Self-Improvement

Extrait original

KernelBench : evaluate correctness and speed for generated GPU kernels. 250 PyTorch tasks to evaluate whether LLM can write fast and correct kernels.
Contexte

The evaluation metric fast_p = the percentage of generated kernels that are correct and faster than baseline.

Lilian Weng
Partager un aperçu

Human–AI performance tradeoffs in RE-Bench

Voir ce sujet

AI agents outperform humans at short time budgets but not longer ones

In RE-Bench, the best AI agents scored 4× higher than human experts at a 2-hour budget, yet humans demonstrated better returns to extended effort and surpassed agents at both 8-hour and 32-hour time budgets.

Éléments favorables

Harness Engineering for Self-Improvement

Extrait original

Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.
Contexte

RE-Bench : evaluate frontier AI agents on realistic ML research-engineering envs against human experts. 7 challenging, open-ended ML research-engineering environments. Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs. Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc. Includes data from 71 eight-hour attempts by 61 distinct human experts. Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions.

Lilian Weng
Partager un aperçu

Propos par date de source1

Les propos sont classés par date de publication de la source originale ; une différence de formulation ne prouve pas un changement de position.