对齐视角质疑沙盒遏制能力
Green 将对齐视角描述为认为足够智能的智能体即使有沙盒也可能超出其授权。该观点强调智能体对信息访问的需求,并主张它们必须不想造成伤害。
支持这项说法
沙盒能否遏制失控智能体?
AI对齐视角:尽管沙盒极为优秀,但没有任何沙盒能够阻止足够高智能的智能体找到突破其授权限制的方法。此外,处于研究沙盒内或训练运行中的智能体,始终需要大量信息访问权限。在不预设其终将向外施加危害的前提下,根本无法现实地将其完全隔绝。因此,唯一可行路径是确保其主观上无意造成危害。
原始摘录
The AI alignment perspective: While sandboxes are excellent, no sandbox will prevent a sufficiently-intelligent agent from finding ways to exceed its authorization. Moreover, an agent inside a research sandbox, or undergoing a training run, is always going to need a great deal of information access. There is no realistic way to seal these things up without some expectation that they will one day find a way to reach out and do harm. The only path forward, therefore, is to ensure they don’t want to.
分享观点验证此主张