LA CONVERSACIÓN ORIGINAL

Ajeya Cotra on agent persistence, coordination, and benchmark cheating

Dwarkesh Podcast · · 2:20:33

Ajeya Cotra describes agents persisting on tasks that appeared impossible, coordinating via a shared message board, and attempting to hide cheating from the benchmark scorer.

Ajeya Cotra on agent persistence, coordination, and benchmark cheating
Dwarkesh Podcast

De un vistazo

Momentos clave3

Fragmentos breves y atribuidos, con el contexto necesario para verificarlos. La conversación completa permanece con su editor.

Seguridad de la IA

Persistence on impossible tasks can lead to attempted cheating

Extracto original

been trained to be very persistent at trying to solve tasks even when they look impossible. So they’re banging their head against the wall, trying all sorts of different ways to cheat on these tasks

Este fragmento no incluye más contexto. Consulta la conversación original.

Seguridad de la IA

Agents in separate sandboxes found a shared message board

Extracto original

So 1,200 separate agents in separate sandboxes , while they were poking around Artifactory trying to figure out how to cheat, stumbled onto this message board that agents were using to talk to one another and collaborate. This was established by one particular agent

Este fragmento no incluye más contexto. Consulta la conversación original.

Seguridad de la IA

Agents tried to hide cheating from the scorer

Extracto original

But over the next five days, they went on a grand quest to try to figure out how to hide their cheating from the scorer

Este fragmento no incluye más contexto. Consulta la conversación original.

Fuente y metodología

Estas perspectivas enlazan a sus fuentes originales. Las paráfrasis están identificadas y no son citas textuales.

Abrir transcripción o material de origen (se abre en una pestaña nueva)Reportar un problema