话题 / AI部署架构

观点转述

Harness 在原始智能之外编排模型执行

Harness 是围绕基础模型的系统,负责编排执行,并决定模型如何思考与规划、调用工具并采取行动、感知与管理上下文、存储产物以及评估结果。

观点背后的信息

译文仅辅助阅读;核查观点请以原始摘录为准。

用于自我改进的 Harness 工程

‘运行框架’(Harness)指围绕基础模型构建的系统,负责编排执行流程,并决定模型如何思考与规划、调用工具与执行动作、感知与管理上下文、存储中间产物,以及评估结果。

原始摘录
A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results.
上下文

我特意使用‘部署系统’这一表述,是因为基础模型与真实世界环境之间的中间层,其重要性不亚于模型本身的原始智能(即预训练后立即开展的评测)。运行框架是 AI 部署中的关键组件,这一点已由 Claude Code 和 Codex 等成功的编程智能体产品所印证。

原始上下文

I explicitly mention “deployment system” because the layer between the raw model and the real-world context seems to be as important as the model’s raw intelligence (i.e. the evals right after pretraining). Harnesses are important components of AI deployment, as shown by successful coding agent products such as Claude Code and Codex.