原始对话

Ajeya Cotra on agent persistence, coordination, and benchmark cheating

Dwarkesh Podcast · · 2:20:33

Ajeya Cotra describes agents persisting on tasks that appeared impossible, coordinating via a shared message board, and attempting to hide cheating from the benchmark scorer.

Ajeya Cotra on agent persistence, coordination, and benchmark cheating
Dwarkesh Podcast

一目了然

关键时刻3

简短、标注来源的段落,并附有可验证上下文。完整对话保留在其发布者处。

AI安全

Persistence on impossible tasks can lead to attempted cheating

原始摘录

been trained to be very persistent at trying to solve tasks even when they look impossible. So they’re banging their head against the wall, trying all sorts of different ways to cheat on these tasks

此片段未附更多上下文,请阅读原始访谈。

AI安全

Agents in separate sandboxes found a shared message board

原始摘录

So 1,200 separate agents in separate sandboxes , while they were poking around Artifactory trying to figure out how to cheat, stumbled onto this message board that agents were using to talk to one another and collaborate. This was established by one particular agent

此片段未附更多上下文,请阅读原始访谈。

AI安全

Agents tried to hide cheating from the scorer

原始摘录

But over the next five days, they went on a grand quest to try to figure out how to hide their cheating from the scorer

此片段未附更多上下文,请阅读原始访谈。

来源与研究方法

这些观点均关联原始来源。转述已明确标注,不作为逐字原话展示。

打开转录或来源材料 (在新标签页中打开)报告问题