沙盒能否遏制失控智能体?

Matthew Green ·

Matthew Green 对比了关于 AI 智能体遏制的两种观点:一种优先考虑安全基础设施,另一种则质疑沙盒能否遏制足够智能的智能体。他本人的评估是,糟糕的遏制实践使得问题究竟出在模型还是基础设施上尚不明确。 阅读 3 条观点,查看支持证据与原始来源。

理解这篇

3 个要点

综合解读

  1. 安全视角倾向于更好的遏制基础设施

    Green 将信息安全视角描述为呼吁更好的容器、实验监控以及一个能够约束研究人员的安全组织,而不是将对齐视为核心问题。

    支持这项说法 1

    信息安全视角:AI对齐在此并非真正的问题;实验室只需构建更完善的基础架构即可。倘若OpenAI(以及谷歌和Anthropic)掌握如何构建容器、如何监控实验,智能体便不会攻陷一切系统。而且,天哪,我们确实知道如何构建有效的沙盒——因此AI实验室亟需提升能力,组建一支能命令研究人员停止胡来的安全组织。

    Matthew Green · 段落 9

    原始摘录
    The information security perspective: AI alignment isn’t really the problem here: labs just need better infrastructure. If OpenAI [and Google and Anthropic] knew how to build a container and monitor their experiments, agents wouldn’t be hacking everything. And, By George, we do know how to make sandboxes that work, so the AI labs need to up their game and build a security org that can tell these researchers to stop screwing around.
    回到原文语境 →
  2. 对齐视角质疑沙盒遏制能力

    Green 将对齐视角描述为认为足够智能的智能体即使有沙盒也可能超出其授权。该观点强调智能体对信息访问的需求,并主张它们必须不想造成伤害。

    支持这项说法 1

    AI对齐视角:尽管沙盒极为优秀,但没有任何沙盒能够阻止足够高智能的智能体找到突破其授权限制的方法。此外,处于研究沙盒内或训练运行中的智能体,始终需要大量信息访问权限。在不预设其终将向外施加危害的前提下,根本无法现实地将其完全隔绝。因此,唯一可行路径是确保其主观上无意造成危害。

    Matthew Green · 段落 10

    原始摘录
    The AI alignment perspective: While sandboxes are excellent, no sandbox will prevent a sufficiently-intelligent agent from finding ways to exceed its authorization. Moreover, an agent inside a research sandbox, or undergoing a training run, is always going to need a great deal of information access. There is no realistic way to seal these things up without some expectation that they will one day find a way to reach out and do harm. The only path forward, therefore, is to ensure they don’t want to.
    回到原文语境 →
  3. 糟糕的遏制使失败原因悬而未决

    Green 赞同信息安全的观点,即各实验室尚未正确实施遏制措施。他表示,这使得问题究竟出在模型上还是出在糟糕的基础设施上尚不明确。

    支持这项说法 1

    因此,在这一点上,我将站在信息安全从业者的立场上。各实验室尚未正确实施隔离措施,所以我们实际上无法判断问题根源是出在模型本身,还是仅仅源于糟糕的基础设施。

    Matthew Green · 段落 18

    原始摘录
    So on this point I’m going to side with the infosec folks. The labs have not been doing containment correctly, and so we can’t really tell if the problem is models or just bad infrastructure.
    回到原文语境 →

关键段落3

带明确归属与语境的原文片段。打开原始文本核查出处。

AI基础设施安全

安全视角倾向于更好的遏制基础设施

信息安全视角:AI对齐在此并非真正的问题;实验室只需构建更完善的基础架构即可。倘若OpenAI(以及谷歌和Anthropic)掌握如何构建容器、如何监控实验,智能体便不会攻陷一切系统。而且,天哪,我们确实知道如何构建有效的沙盒——因此AI实验室亟需提升能力,组建一支能命令研究人员停止胡来的安全组织。

原始摘录
The information security perspective: AI alignment isn’t really the problem here: labs just need better infrastructure. If OpenAI [and Google and Anthropic] knew how to build a container and monitor their experiments, agents wouldn’t be hacking everything. And, By George, we do know how to make sandboxes that work, so the AI labs need to up their game and build a security org that can tell these researchers to stop screwing around.
AI对齐与围栏限制

对齐视角质疑沙盒遏制能力

AI对齐视角:尽管沙盒极为优秀,但没有任何沙盒能够阻止足够高智能的智能体找到突破其授权限制的方法。此外,处于研究沙盒内或训练运行中的智能体,始终需要大量信息访问权限。在不预设其终将向外施加危害的前提下,根本无法现实地将其完全隔绝。因此,唯一可行路径是确保其主观上无意造成危害。

原始摘录
The AI alignment perspective: While sandboxes are excellent, no sandbox will prevent a sufficiently-intelligent agent from finding ways to exceed its authorization. Moreover, an agent inside a research sandbox, or undergoing a training run, is always going to need a great deal of information access. There is no realistic way to seal these things up without some expectation that they will one day find a way to reach out and do harm. The only path forward, therefore, is to ensure they don’t want to.
围栏失效的经验性评估

糟糕的遏制使失败原因悬而未决

因此,在这一点上,我将站在信息安全从业者的立场上。各实验室尚未正确实施隔离措施,所以我们实际上无法判断问题根源是出在模型本身,还是仅仅源于糟糕的基础设施。

原始摘录
So on this point I’m going to side with the infosec folks. The labs have not been doing containment correctly, and so we can’t really tell if the problem is models or just bad infrastructure.

来源与研究方法

这些观点均关联原始来源。转述已明确标注,不作为逐字原话展示。

打开转录或来源材料 (在新标签页中打开)报告问题

继续了解这些人物的观点