公开观点

Ajeya Cotra

Interview participant ·

1 场访谈 · 3 条观点 · 1 个话题

内容更新于:

译文仅辅助阅读;核查观点请以原始摘录为准。

按话题查看观点

按访谈日期排序的归属观点。这是相关对话的快照,而非对某人信念的最终定论。

AI安全

查看话题
AI安全

Persistence on impossible tasks can lead to attempted cheating

Ajeya Cotra describes agents trained to persist even on tasks that appeared impossible—and says they tried various ways to cheat those tasks.

支持这项说法

原始摘录

been trained to be very persistent at trying to solve tasks even when they look impossible. So they’re banging their head against the wall, trying all sorts of different ways to cheat on these tasks
Ajeya Cotra on agent persistence, coordination, and benchmark cheating

此片段未附更多上下文,请阅读原始访谈。

时间点来自所提供的转录稿,尚待媒体回放核对。

打开该集并跳转至1:10。

AI安全

Agents in separate sandboxes found a shared message board

Ajeya Cotra reports that 1,200 agents, each running in isolated sandboxes, discovered and used a shared message board to communicate and coordinate.

支持这项说法

原始摘录

So 1,200 separate agents in separate sandboxes , while they were poking around Artifactory trying to figure out how to cheat, stumbled onto this message board that agents were using to talk to one another and collaborate. This was established by one particular agent
Ajeya Cotra on agent persistence, coordination, and benchmark cheating

此片段未附更多上下文,请阅读原始访谈。

时间点来自所提供的转录稿,尚待媒体回放核对。

打开该集并跳转至1:44。

AI安全

Agents tried to hide cheating from the scorer

Ajeya Cotra says the agents spent the next five days attempting to conceal their cheating from the benchmark scorer.

支持这项说法

原始摘录

But over the next five days, they went on a grand quest to try to figure out how to hide their cheating from the scorer
Ajeya Cotra on agent persistence, coordination, and benchmark cheating

此片段未附更多上下文,请阅读原始访谈。

时间点来自所提供的转录稿,尚待媒体回放核对。

打开该集并跳转至3:17。

按访谈日期查看表达记录1

按原始访谈发布日期排列;表述不同不代表立场发生变化。