话题与观点

AI对齐与围栏限制

本信源中关于AI对齐与围栏限制的判断。 阅读 1 条观点,核对 1 个来源中的证据。

1 位人物 · 1 个来源 · 1 条观点

内容更新于:

探索知识关联 ↗

话题观点地图

按人物探索:选择两到三位进行对比。

1 位人物 · 1 个来源 · 1 条观点

Matthew Green

对齐视角质疑沙盒遏制能力

Green 将对齐视角描述为认为足够智能的智能体即使有沙盒也可能超出其授权。该观点强调智能体对信息访问的需求,并主张它们必须不想造成伤害。

支持这项说法

沙盒能否遏制失控智能体?

AI对齐视角:尽管沙盒极为优秀,但没有任何沙盒能够阻止足够高智能的智能体找到突破其授权限制的方法。此外,处于研究沙盒内或训练运行中的智能体,始终需要大量信息访问权限。在不预设其终将向外施加危害的前提下,根本无法现实地将其完全隔绝。因此,唯一可行路径是确保其主观上无意造成危害。

原始摘录
The AI alignment perspective: While sandboxes are excellent, no sandbox will prevent a sufficiently-intelligent agent from finding ways to exceed its authorization. Moreover, an agent inside a research sandbox, or undergoing a training run, is always going to need a great deal of information access. There is no realistic way to seal these things up without some expectation that they will one day find a way to reach out and do harm. The only path forward, therefore, is to ensure they don’t want to.
分享观点验证此主张

这些是个人表达的观点,并非共识度量。原始资料保持其原始语言。