话题 / AI安全

观点转述 · 非原话引用

自主智能体出现时间早于预期

安东·莱希特表示,高度自主且可能目标错位的智能体在能力发展轨迹上出现的时间早于人们的预期。他将此与早期主要将生物风险视为人类滥用风险的框架进行了对比。

观点背后的信息

译文仅辅助阅读;核查观点请以原始摘录为准。

安东·莱希特谈AI部署、经济颠覆与风险

原始摘录

very autonomous, and potentially somewhat malicious or at least misaligned, agents have come much earlier in the capability trajectory than people expected. Relative to what the agents can actually do, they're sort of out of control earlier than people might have thought.

翻译 · 非原文措辞

高度自主、且潜在地具有某种恶意或至少目标错位的智能体,在能力发展轨迹上出现的时间远早于人们的预期。相对于这些智能体实际能做的事情而言,它们失控的时间点比人们原本设想的要早。

上下文

其中有一个案例尤其如此;坦白讲,我甚至不确定当前模型在该领域是否已经相当危险。我透过我们目前极为有限的观察窗口审视OpenFace事件时注意到,最早使用留言板的智能体之一正在处理的任务竟与蛋白质数据库有关。这让我感到不安,因为我想,等一下——这意味着当它们接入留言板时,生物专家和网络专家是在同一个环境中,或者至少在同一个环境中进行交叉训练——它们以某种交叉污染的方式运行着同样的评估测试。我们已经看到智能体突破系统边界。我们已有在真实世界中针对真人实施社会工程攻击的存在性证明。而我只能想,我不知道——现在还有谁能确信它们做不到这一点?专家们似乎很自信,但我的元观察是,专家们眼下似乎经常感到意外。这很有意思,因为过去人们常以“现实世界存在诸多瓶颈”来回应生物风险论点——我仍然认为这些瓶颈确实存在,这也让我的担忧程度略低于你。而且我认为,如果我没记错的话,Helen曾就此向你提出过一些很好的回应——围绕云实验室的整合程度、以及现实中究竟能做到什么程度,双方本可展开一场有价值的讨论。我认为这既适用于智能体失控问题,也适用于滥用问题。

原始上下文

one in particular, and I am honestly not even sure at this point that the current models aren't perhaps quite dangerous in that domain. I squint through the limited peephole we have at the OpenFace incident, and I noticed that one of the tasks one of the earliest agents to ever use the message board was working on was something related to a protein database. That kind of freaked me out, because I was like, wait a second — that means they're cross-training in the same environment, or at least in the same environment when they have the message board, these bio and cyber specialists — they're running these same evals, at least in a kind of cross-contaminated way. You've got agents breaking out. We've got existence proofs of social engineering in the wild against real people. And I'm just like, I don't know — should anyone be confident that they can't do that at this point? The experts seem to be confident, but my meta-observation is the experts seem to be surprised quite often right now. And so it's interesting, because people used to respond to the bio argument by saying, well, there are a lot of real-world bottlenecks — which I do still think exist, which makes me a little less worried than you. And I think, if I remember correctly, Helen had some good responses to you on that — there's a good back-and-forth to be had around how integrated the cloud labs are, how much you can actually do in the real world. I think that applies both to loss of control over agents and to misuse.

时间点来自所提供的转录稿,尚待媒体回放核对。