话题与观点

AI系统安全与可观测性

本信息来源中关于AI系统安全与可观测性的判断。 阅读 1 条观点,核对 1 个来源中的证据。

1 位人物 · 1 个来源 · 1 条观点

内容更新于:

探索知识关联 ↗

话题观点地图

按人物探索:选择两到三位进行对比。

1 位人物 · 1 个来源 · 1 条观点

Lilian Weng

Harness 修改必须限定在可编辑界面上,并设置只读保护机制

Harness 修改仅限于 harness 工作区范围内;runs 目录、追踪器(tracer)、验证器(verifier)及大语言模型(LLM)配置均为只读,以防止奖励作弊行为——例如禁用验证器、更换模型或提高推理预算——从而确保所有性能提升均唯一归因于 harness 的修改。

支持这项说法

用于自我改进的 Harness 工程

所有编辑仅作用于运行框架工作区(harness workspace);runs 目录、tracer(追踪器)、verifier(验证器)及大语言模型配置均为只读。此举禁用了若干类奖励黑客行为(例如禁用验证器、替换模型或提高推理预算),确保所有可归因的性能提升均源自运行框架本身的编辑。

原始摘录
Edits are only applied to the harness workspace. the runs directory, tracer, verifier, and LLM configuration are read-only, which disables a set of reward hacking (e.g disabling the verifier, swapping the model, or raising the reasoning budget) and thus it can keep every recorded gain attributable to harness edits.
上下文

决策可观测性(Decision observability):每次编辑均需附带对下一轮效果的预测,以供验证。一个智能体(‘Evolve agent’)读取代码仓库,判断需修改的组件,继而生成具体编辑内容及其推理依据。每次编辑均为文件级、可证伪的主张,并可在下一轮中验证,且须满足两项约束:(1)……(2)编辑须以证据为驱动,并附有明确声明条目:失败证据名称、推断出的根因、拟实施的修复方案,以及预测影响——包括预期修复效果及潜在引发的回归风险。

原始上下文

Decision observability : every edit is paired with a prediction for the next round to validate. An agent (“Evolve agent”) reads the repo and decides which component to edit, and then produces the edit and the reasoning behind it. Every edit is a file-level, falsifiable claim and can be verified in the next round, under two constraints: (1) (2) Edits are evidence-driven, with a manifesto entry: the failure evidence’s name, the inferred root cause, the targeted fix, and a predicted impact comprising both expected fixes and at-risk regressions.

分享观点验证此主张

这些是个人表达的观点,并非共识度量。原始资料保持其原始语言。