PENSÉES EN RELATION
Atlas des connaissances
Explorez les personnes, leurs points de vue et les sources originales.
1 personnes · 1 sources · 10 opinions exprimées
EN CONTEXTE
Lilian Weng
Choisissez un point de vue et retrouvez la conversation originale.
Harness edits must be limited to editable surfaces, with read-only safeguards
Les secteurs égaux servent de repères, pas de classement.
Points de vue 1–4 sur 10 · Sources les plus récentes en premier
Point de vue sélectionné
Harness edits must be limited to editable surfaces, with read-only safeguards
Harness edits are restricted to the harness workspace only; the runs directory, tracer, verifier, and LLM configuration are read-only to prevent reward hacking—such as disabling the verifier, swapping the model, or raising the reasoning budget—ensuring gains remain attributable solely to harness changes.
Il s’agit de points de vue individuels, non d’une mesure du consensus. Le matériel source reste dans sa langue d’origine.
Éléments favorables
Extrait original
Edits are only applied to the harness workspace. the runs directory, tracer, verifier, and LLM configuration are read-only, which disables a set of reward hacking (e.g disabling the verifier, swapping the model, or raising the reasoning budget) and thus it can keep every recorded gain attributable to harness edits.
Contexte
Decision observability : every edit is paired with a prediction for the next round to validate. An agent (“Evolve agent”) reads the repo and decides which component to edit, and then produces the edit and the reasoning behind it. Every edit is a file-level, falsifiable claim and can be verified in the next round, under two constraints: (1) (2) Edits are evidence-driven, with a manifesto entry: the failure evidence’s name, the inferred root cause, the targeted fix, and a predicted impact comprising both expected fixes and at-risk regressions.
Les dates concernent les sources, pas des changements d’opinion. Les textes sans traduction révisée restent dans leur langue originale.