OPINIONES PÚBLICAS

Lilian Weng

1 fuentes · 10 opiniones · 10 temas

Contenido actualizado:

Lilian Weng sobre agent system design, AI deployment architecture, AI evaluation and generalization. Explora 10 puntos de vista por tema, con evidencias de 1 fuente.

Explorar conexiones

Perspectivas por tema

Opiniones atribuidas, ordenadas por fecha de publicación de la fuente. Una muestra de estas intervenciones, no una definición completa de las creencias de la persona.

Las traducciones son para facilitar la lectura; los extractos originales siguen siendo la source evidence.

AI deployment architecture

Ver este tema

Harnesses orchestrate model execution beyond raw intelligence

A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results.

Evidencia a favor

Harness Engineering for Self-Improvement

Extracto original

A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results.
Contexto

I explicitly mention “deployment system” because the layer between the raw model and the real-world context seems to be as important as the model’s raw intelligence (i.e. the evals right after pretraining). Harnesses are important components of AI deployment, as shown by successful coding agent products such as Claude Code and Codex.

Lilian Weng
Compartir información

agent system design

Ver este tema

Harness engineering extends beyond prompt templates to runtime design

Compared with early agent frameworks, harness engineering additionally includes workflow design (e.g. loop engineering), evaluation, permission controls, and persistent state management — it is closer to runtime and software system design: how the model observes, acts, memorizes, checks itself, and improves.

Evidencia a favor

Harness Engineering for Self-Improvement

Extracto original

Compared with early agent frameworks , “agent = LLM + memory + tools + planning + action”, harnesses engineering additionally include workflow design (e.g. loop engineering), evaluation, permission controls, and persistent state management . It is no longer only prompt templates, but closer to runtime and software system design: how the model observes, acts, memorizes, checks itself, and improves.
Lilian Weng
Compartir información

workflow automation

Ver este tema

Goal-oriented loops let models test and iterate on their own

Defining a workflow where the model can operate, test, and iterate is central to automation. A common pattern is a goal-oriented loop: plan → execute → observe/test → improve → execute again until the goal is achieved. The loop may proactively ask users for clarification on task specification or execution preferences.

Evidencia a favor

Harness Engineering for Self-Improvement

Extracto original

Defining a workflow in which the model can operate, test, and iterate is a key design for automation. Karpathy’s autoresearch repo ( https://github.com/karpathy/autoresearch ) is a clean example of how such a workflow can be constructed. A common workflow follows a goal-oriented loop of plan, execute, observe/test, improve, and execute again until the goal is achieved. The process may trigger proactive requests to users for clarity in task specification or execution preference.
Lilian Weng
Compartir información

failure analysis in AI systems

Ver este tema

Root-cause failure analysis requires detailed trace records

To uncover root causes of failures, failure records must contain the terminal verifier-level cause, the causal status of relevant agent behavior, and the abstract agent mechanism exposed by the trace—because surface-level verifier outcomes (e.g., timeout or missing artifact) can mask distinct underlying causal mechanisms.

Evidencia a favor

Harness Engineering for Self-Improvement

Extracto original

Weakness mining : cluster failures into verifier-grounded failure patterns. The current harness $h_t$ is used to evaluate on tasks and execution traces are collected for analysis. Note that two runs can share the same verifier outcome in the error logs on the surface, such as timeout or missing artifact, while having different causal mechanisms. Therefore we need a failure record of rich information, containing the terminal verifier-level cause, the causal status of the relevant agent behavior, and the abstract agent mechanism exposed by the trace, to uncover the root causes.
Lilian Weng
Compartir información

AI system security and observability

Ver este tema

Harness edits must be limited to editable surfaces, with read-only safeguards

Harness edits are restricted to the harness workspace only; the runs directory, tracer, verifier, and LLM configuration are read-only to prevent reward hacking—such as disabling the verifier, swapping the model, or raising the reasoning budget—ensuring gains remain attributable solely to harness changes.

Evidencia a favor

Harness Engineering for Self-Improvement

Extracto original

Edits are only applied to the harness workspace. the runs directory, tracer, verifier, and LLM configuration are read-only, which disables a set of reward hacking (e.g disabling the verifier, swapping the model, or raising the reasoning budget) and thus it can keep every recorded gain attributable to harness edits.
Contexto

Decision observability : every edit is paired with a prediction for the next round to validate. An agent (“Evolve agent”) reads the repo and decides which component to edit, and then produces the edit and the reasoning behind it. Every edit is a file-level, falsifiable claim and can be verified in the next round, under two constraints: (1) (2) Edits are evidence-driven, with a manifesto entry: the failure evidence’s name, the inferred root cause, the targeted fix, and a predicted impact comprising both expected fixes and at-risk regressions.

Lilian Weng
Compartir información

AI evaluation and generalization

Ver este tema

Evolved harnesses generalize beyond their original benchmarks

A harness evolved on Terminal-Bench-2 transferred without further evolution to SWE-bench-verified, indicating it encodes general engineering experience in harness components rather than benchmark-specific optimization.

Evidencia a favor

Harness Engineering for Self-Improvement

Extracto original

On Terminal-Bench-2, AHE achieved better than human-designed harness (OpenCode, Terminus-2, Codex) except for Hard tier and a few other self-evolve baselines (ACE, TF-GRPO). The same frozen harness, without further evolving, transfers to SWE-bench-verified, indicating that the evolved harness is able to encode engineering experience into harness components rather than doing benchmark-specific optimization.
Lilian Weng
Compartir información

RE-Bench AI agent evaluation methodology

Ver este tema

RE-Bench evaluates AI agents on realistic ML engineering tasks, using human performance as a benchmark

RE-Bench evaluates frontier AI agents on 7 open-ended ML research-engineering environments, each defined by a scoring function, starting solution, and reference solution, runnable with 8 or fewer H100 GPUs. It includes data from 71 eight-hour attempts by 61 distinct human experts, where humans achieved non-zero scores in 82% of attempts and matched or exceeded strong reference solutions in 24%.

Evidencia a favor

Harness Engineering for Self-Improvement

Extracto original

RE-Bench : evaluate frontier AI agents on realistic ML research-engineering envs against human experts. 7 challenging, open-ended ML research-engineering environments.
Contexto

Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs. Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc. Includes data from 71 eight-hour attempts by 61 distinct human experts. Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions. Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.

Lilian Weng
Compartir información

MLE-bench offline Kaggle benchmark

Ver este tema

MLE-bench uses 75 Kaggle competitions as offline ML engineering benchmarks

MLE-bench evaluates ML engineering agents on 75 curated offline Kaggle competitions, testing model training, dataset preparation, experiment execution, and submission to grading scripts. Kaggle public leaderboards serve as human baselines. The best-performing setup—o1-preview with AIDE scaffolding—reached at least Kaggle bronze-medal level in 16.9% of competitions.

Evidencia a favor

Harness Engineering for Self-Improvement

Extracto original

MLE-bench : evaluate ML engineering agents on offline Kaggle competitions. Contains 75 ML-engineering competitions curated from Kaggle.
Contexto

Tests training models, preparing datasets, running experiments, and submitting predictions to grading scripts. Uses Kaggle public leaderboards as human baselines. Best setup in the paper, o1-preview with AIDE scaffolding, reached at least Kaggle bronze-medal level in 16.9% of competitions. Includes resource-scaling and contamination analyses.

Lilian Weng
Compartir información

KernelBench GPU kernel evaluation

Ver este tema

KernelBench measures correctness and speed of generated GPU kernels

KernelBench evaluates LLMs on 250 PyTorch tasks to assess their ability to generate correct and fast GPU kernels, using fast_p—the percentage of generated kernels that are both correct and faster than the baseline—as its metric.

Evidencia a favor

Harness Engineering for Self-Improvement

Extracto original

KernelBench : evaluate correctness and speed for generated GPU kernels. 250 PyTorch tasks to evaluate whether LLM can write fast and correct kernels.
Contexto

The evaluation metric fast_p = the percentage of generated kernels that are correct and faster than baseline.

Lilian Weng
Compartir información

Human–AI performance tradeoffs in RE-Bench

Ver este tema

AI agents outperform humans at short time budgets but not longer ones

In RE-Bench, the best AI agents scored 4× higher than human experts at a 2-hour budget, yet humans demonstrated better returns to extended effort and surpassed agents at both 8-hour and 32-hour time budgets.

Evidencia a favor

Harness Engineering for Self-Improvement

Extracto original

Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.
Contexto

RE-Bench : evaluate frontier AI agents on realistic ML research-engineering envs against human experts. 7 challenging, open-ended ML research-engineering environments. Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs. Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc. Includes data from 71 eight-hour attempts by 61 distinct human experts. Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions.

Lilian Weng
Compartir información

Declaraciones por fecha de la fuente1

Las declaraciones se ordenan por fecha de publicación de la fuente original; las diferencias de redacción no demuestran un cambio de postura.