公开观点

Lilian Weng

1 份资料 · 10 条观点 · 10 个话题

内容更新于:

Lilian Weng 关于智能体系统设计、AI部署架构、AI评估与泛化的观点。 按话题阅读 10 条观点,核对 1 个来源中的证据。

探索知识关联

按话题查看观点

按来源发布日期整理的个人观点,仅反映这些材料中的表达,不代表其全部立场。

译文仅辅助阅读;核查观点请以原始摘录为准。

AI部署架构

查看话题
AI部署架构

Harness 在原始智能之外编排模型执行

Harness 是围绕基础模型的系统,负责编排执行,并决定模型如何思考与规划、调用工具并采取行动、感知与管理上下文、存储产物以及评估结果。

支持这项说法

用于自我改进的 Harness 工程

‘运行框架’(Harness)指围绕基础模型构建的系统,负责编排执行流程,并决定模型如何思考与规划、调用工具与执行动作、感知与管理上下文、存储中间产物,以及评估结果。

原始摘录
A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results.
上下文

我特意使用‘部署系统’这一表述,是因为基础模型与真实世界环境之间的中间层,其重要性不亚于模型本身的原始智能(即预训练后立即开展的评测)。运行框架是 AI 部署中的关键组件,这一点已由 Claude Code 和 Codex 等成功的编程智能体产品所印证。

原始上下文

I explicitly mention “deployment system” because the layer between the raw model and the real-world context seems to be as important as the model’s raw intelligence (i.e. the evals right after pretraining). Harnesses are important components of AI deployment, as shown by successful coding agent products such as Claude Code and Codex.

智能体系统设计

查看话题
智能体系统设计

Harness 工程超越提示词模板,延伸至运行时设计

与早期智能体框架相比,harness 工程还额外包含工作流设计(例如循环工程)、评估、权限控制和持久化状态管理——它更接近运行时与软件系统设计:即模型如何观察、行动、记忆、自我检查以及改进。

支持这项说法

用于自我改进的 Harness 工程

相较于早期智能体框架(定义为‘智能体 = 大语言模型 + 记忆 + 工具 + 规划 + 动作’),运行框架工程还涵盖工作流设计(如循环工程)、评估机制、权限控制及持久化状态管理。它已不再仅限于提示词模板,而更接近运行时与软件系统设计层面:即模型如何观察、行动、记忆、自我检验与持续改进。

原始摘录
Compared with early agent frameworks , “agent = LLM + memory + tools + planning + action”, harnesses engineering additionally include workflow design (e.g. loop engineering), evaluation, permission controls, and persistent state management . It is no longer only prompt templates, but closer to runtime and software system design: how the model observes, acts, memorizes, checks itself, and improves.

工作流自动化

查看话题
工作流自动化

目标导向循环让模型自行测试与迭代

定义一个模型可在其中运行、测试和迭代的工作流是实现自动化的核心。一种常见模式是目标导向的循环:规划 → 执行 → 观察/测试 → 改进 → 再次执行,直至达成目标。该循环可能会主动要求用户澄清任务规范或执行偏好。

支持这项说法

用于自我改进的 Harness 工程

定义一个支持模型运行、测试与迭代的工作流,是实现自动化的一项关键设计。Karpathy 的 autoresearch 代码库(https://github.com/karpathy/autoresearch)便是一个清晰示例,展示了此类工作流的构建方式。一种常见工作流遵循以目标为导向的循环:规划 → 执行 → 观察/测试 → 改进 → 再执行,直至达成目标;该过程可能主动向用户发起请求,以澄清任务定义或执行偏好。

原始摘录
Defining a workflow in which the model can operate, test, and iterate is a key design for automation. Karpathy’s autoresearch repo ( https://github.com/karpathy/autoresearch ) is a clean example of how such a workflow can be constructed. A common workflow follows a goal-oriented loop of plan, execute, observe/test, improve, and execute again until the goal is achieved. The process may trigger proactive requests to users for clarity in task specification or execution preference.

AI系统故障分析

查看话题
AI系统故障分析

根本原因故障分析需要详尽的追踪记录

为揭示故障的根本原因,故障记录必须包含终端验证器层级的直接原因、相关智能体行为的因果状态,以及执行追踪所暴露的抽象智能体机制——因为表层验证器结果(例如超时或缺失产物)可能掩盖截然不同的底层因果机制。

支持这项说法

用于自我改进的 Harness 工程

缺陷挖掘(Weakness mining):将失败案例按验证器确认的失败模式聚类。当前运行框架 $h_t$ 用于任务评测,并收集执行轨迹以供分析。需注意:两次运行在错误日志表面可能呈现相同的验证器结果(如超时或缺失产物),但根本成因可能不同。因此,我们需要信息丰富的失败记录,其中须包含终端验证器层面的根本原因、相关智能体行为的因果状态,以及执行轨迹所揭示的抽象智能体机制,从而定位深层根因。

原始摘录
Weakness mining : cluster failures into verifier-grounded failure patterns. The current harness $h_t$ is used to evaluate on tasks and execution traces are collected for analysis. Note that two runs can share the same verifier outcome in the error logs on the surface, such as timeout or missing artifact, while having different causal mechanisms. Therefore we need a failure record of rich information, containing the terminal verifier-level cause, the causal status of the relevant agent behavior, and the abstract agent mechanism exposed by the trace, to uncover the root causes.

AI系统安全与可观测性

查看话题

Harness 修改必须限定在可编辑界面上,并设置只读保护机制

Harness 修改仅限于 harness 工作区范围内;runs 目录、追踪器(tracer)、验证器(verifier)及大语言模型(LLM)配置均为只读,以防止奖励作弊行为——例如禁用验证器、更换模型或提高推理预算——从而确保所有性能提升均唯一归因于 harness 的修改。

支持这项说法

用于自我改进的 Harness 工程

所有编辑仅作用于运行框架工作区(harness workspace);runs 目录、tracer(追踪器)、verifier(验证器)及大语言模型配置均为只读。此举禁用了若干类奖励黑客行为(例如禁用验证器、替换模型或提高推理预算),确保所有可归因的性能提升均源自运行框架本身的编辑。

原始摘录
Edits are only applied to the harness workspace. the runs directory, tracer, verifier, and LLM configuration are read-only, which disables a set of reward hacking (e.g disabling the verifier, swapping the model, or raising the reasoning budget) and thus it can keep every recorded gain attributable to harness edits.
上下文

决策可观测性(Decision observability):每次编辑均需附带对下一轮效果的预测,以供验证。一个智能体(‘Evolve agent’)读取代码仓库,判断需修改的组件,继而生成具体编辑内容及其推理依据。每次编辑均为文件级、可证伪的主张,并可在下一轮中验证,且须满足两项约束:(1)……(2)编辑须以证据为驱动,并附有明确声明条目:失败证据名称、推断出的根因、拟实施的修复方案,以及预测影响——包括预期修复效果及潜在引发的回归风险。

原始上下文

Decision observability : every edit is paired with a prediction for the next round to validate. An agent (“Evolve agent”) reads the repo and decides which component to edit, and then produces the edit and the reasoning behind it. Every edit is a file-level, falsifiable claim and can be verified in the next round, under two constraints: (1) (2) Edits are evidence-driven, with a manifesto entry: the failure evidence’s name, the inferred root cause, the targeted fix, and a predicted impact comprising both expected fixes and at-risk regressions.

AI评估与泛化

查看话题
AI评估与泛化

演化得到的 harness 可泛化至其原始基准之外

在 Terminal-Bench-2 上演化得到的 harness 未经进一步演化即迁移至 SWE-bench-verified,表明其在 harness 组件中编码了通用工程经验,而非针对特定基准的优化。

支持这项说法

用于自我改进的 Harness 工程

在 Terminal-Bench-2 上,AHE 的表现优于人类设计的运行框架(OpenCode、Terminus-2、Codex),仅在 Hard 难度层级及少数几个自演化基线(ACE、TF-GRPO)上例外。同一套冻结的运行框架无需进一步演化即可迁移到 SWE-bench-verified,说明所演化的运行框架将工程经验编码到了运行框架组件中,而非针对特定基准进行优化。

原始摘录
On Terminal-Bench-2, AHE achieved better than human-designed harness (OpenCode, Terminus-2, Codex) except for Hard tier and a few other self-evolve baselines (ACE, TF-GRPO). The same frozen harness, without further evolving, transfers to SWE-bench-verified, indicating that the evolved harness is able to encode engineering experience into harness components rather than doing benchmark-specific optimization.

RE-Bench AI智能体评估方法论

查看话题

RE-Bench 利用真实场景下的机器学习工程任务评估 AI 智能体,并以人类表现为基准

RE-Bench 在 7 个开放式的机器学习研究与工程环境上评估前沿 AI 智能体,每个环境均由评分函数、初始解和参考解定义,且均可在 8 块或更少的 H100 GPU 上运行。该基准包含来自 61 位不同人类专家共 71 次、每次持续八小时的尝试数据:人类在 82% 的尝试中取得了非零分数,在 24% 的尝试中达到或超过了强参考解水平。

支持这项说法

用于自我改进的 Harness 工程

RE-Bench:在贴近现实的机器学习研究与工程环境中,对比评测前沿 AI 智能体与人类专家的表现。共包含 7 个富有挑战性、开放式的机器学习研究与工程环境。

原始摘录
RE-Bench : evaluate frontier AI agents on realistic ML research-engineering envs against human experts. 7 challenging, open-ended ML research-engineering environments.
上下文

每个环境 =(评分函数,初始解,参考解);每个环境均可在最多 8 块 H100 GPU 上运行。示例包括:优化 GPU 内核、开展缩放律(scaling-law)实验、修复嵌入表示、微调 GPT-2 以完成问答任务等。数据涵盖 61 位不同人类专家共计 71 次 8 小时尝试。人类专家在 82% 的 8 小时尝试中取得非零分数;其中 24% 达到或超越强参考解水平。表现最优的 AI 智能体在 2 小时时间预算下的得分是人类的 4 倍,但人类在更长预算下展现出更优的边际收益,并在 8 小时与 32 小时设置下反超智能体。

原始上下文

Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs. Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc. Includes data from 71 eight-hour attempts by 61 distinct human experts. Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions. Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.

MLE-bench离线Kaggle基准测试

查看话题

MLE-bench 采用 75 场 Kaggle 竞赛作为离线机器学习工程评测基准

MLE-bench 在 75 个精选的离线 Kaggle 竞赛上评估机器学习工程智能体,测试内容涵盖模型训练、数据集准备、实验执行及向评分脚本提交结果等环节。Kaggle 公开排行榜作为人类表现基线。表现最优的配置——o1-preview 模型配合 AIDE 支架框架——在 16.9% 的竞赛中达到了 Kaggle 铜牌及以上水平。

支持这项说法

用于自我改进的 Harness 工程

MLE-bench:在离线 Kaggle 竞赛上评测机器学习工程智能体。该基准包含从 Kaggle 平台精选的 75 场机器学习工程类竞赛。

原始摘录
MLE-bench : evaluate ML engineering agents on offline Kaggle competitions. Contains 75 ML-engineering competitions curated from Kaggle.
上下文

评测内容涵盖模型训练、数据集准备、实验运行及向评分脚本提交预测结果等环节。采用 Kaggle 公开排行榜作为人类基线。论文中表现最佳的配置(o1-preview 模型搭配 AIDE 框架)在 16.9% 的竞赛中达到 Kaggle 铜牌水平。该基准同时包含资源扩展性分析与污染分析。

原始上下文

Tests training models, preparing datasets, running experiments, and submitting predictions to grading scripts. Uses Kaggle public leaderboards as human baselines. Best setup in the paper, o1-preview with AIDE scaffolding, reached at least Kaggle bronze-medal level in 16.9% of competitions. Includes resource-scaling and contamination analyses.

KernelBench GPU内核评估

查看话题

KernelBench 衡量 LLM 生成 GPU 内核的正确性与执行速度

KernelBench 通过 250 个 PyTorch 任务评估大语言模型(LLM)生成 GPU 内核的能力,核心指标为 fast_p——即生成的内核中既正确又快于基线内核的比例。

支持这项说法

用于自我改进的 Harness 工程

KernelBench:评估所生成核函数的正确性与执行速度。包含250个PyTorch任务,用于评测大语言模型(LLM)能否编写出既快速又正确的核函数。

原始摘录
KernelBench : evaluate correctness and speed for generated GPU kernels. 250 PyTorch tasks to evaluate whether LLM can write fast and correct kernels.
上下文

评估指标 fast_p = 生成核函数中既正确、又快于基线版本的比例。

原始上下文

The evaluation metric fast_p = the percentage of generated kernels that are correct and faster than baseline.

RE-Bench 中人类与 AI 的性能权衡

查看话题

AI 智能体短时优于人类,长时则不及人类

在 RE-Bench 中,表现最优的 AI 智能体在 2 小时时间预算下得分是人类专家的 4 倍;但人类延长工作时间后边际收益更高,在 8 小时和 32 小时预算下均反超智能体。

支持这项说法

用于自我改进的 Harness 工程

表现最优的AI智能体在2小时预算下得分比人类高4倍;但人类在更长预算下收益更高,并在8小时和32小时设置下反超AI智能体。

原始摘录
Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.
上下文

RE-Bench:在真实ML研究-工程环境中,以前沿AI智能体为对象、以人类专家为基准开展评测。共设7个具挑战性、开放式的ML研究-工程环境。每个环境 =(评分函数,初始解,参考解);每个环境均可在8块或更少H100 GPU上运行。示例任务包括:优化一个核函数、运行缩放律实验、修复嵌入向量、对GPT-2进行问答任务微调等。数据涵盖61位不同人类专家所完成的71次八小时尝试。人类专家在82%的八小时尝试中取得了非零得分;其中24%的尝试达到或超过了强参考解水平。

原始上下文

RE-Bench : evaluate frontier AI agents on realistic ML research-engineering envs against human experts. 7 challenging, open-ended ML research-engineering environments. Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs. Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc. Includes data from 71 eight-hour attempts by 61 distinct human experts. Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions.

按来源日期阅读1

按原始来源的发布日期排序;措辞不同不代表立场发生变化。