Anton Leicht on AI deployment, economic disruption and risk
Original excerpt
very autonomous, and potentially somewhat malicious or at least misaligned, agents have come much earlier in the capability trajectory than people expected. Relative to what the agents can actually do, they're sort of out of control earlier than people might have thought.
Context
one in particular, and I am honestly not even sure at this point that the current models aren't perhaps quite dangerous in that domain. I squint through the limited peephole we have at the OpenFace incident, and I noticed that one of the tasks one of the earliest agents to ever use the message board was working on was something related to a protein database. That kind of freaked me out, because I was like, wait a second — that means they're cross-training in the same environment, or at least in the same environment when they have the message board, these bio and cyber specialists — they're running these same evals, at least in a kind of cross-contaminated way. You've got agents breaking out. We've got existence proofs of social engineering in the wild against real people. And I'm just like, I don't know — should anyone be confident that they can't do that at this point? The experts seem to be confident, but my meta-observation is the experts seem to be surprised quite often right now. And so it's interesting, because people used to respond to the bio argument by saying, well, there are a lot of real-world bottlenecks — which I do still think exist, which makes me a little less worried than you. And I think, if I remember correctly, Helen had some good responses to you on that — there's a good back-and-forth to be had around how integrated the cloud labs are, how much you can actually do in the real world. I think that applies both to loss of control over agents and to misuse.
Start time comes from the supplied transcript. Playback alignment is awaiting review.