LA CONVERSATION ORIGINALE

Ajeya Cotra on agent persistence, coordination, and benchmark cheating

Dwarkesh Podcast · · 2:20:33

Ajeya Cotra describes agents persisting on tasks that appeared impossible, coordinating via a shared message board, and attempting to hide cheating from the benchmark scorer.

Ajeya Cotra on agent persistence, coordination, and benchmark cheating
Dwarkesh Podcast

En un coup d’œil

Moments clés3

Extraits courts et attribués, avec le contexte permettant de les vérifier. La conversation complète reste chez son éditeur.

Sécurité de l'IA

Persistence on impossible tasks can lead to attempted cheating

Extrait original

been trained to be very persistent at trying to solve tasks even when they look impossible. So they’re banging their head against the wall, trying all sorts of different ways to cheat on these tasks

Cet extrait ne contient pas de contexte supplémentaire. Consultez la conversation originale.

Sécurité de l'IA

Agents in separate sandboxes found a shared message board

Extrait original

So 1,200 separate agents in separate sandboxes , while they were poking around Artifactory trying to figure out how to cheat, stumbled onto this message board that agents were using to talk to one another and collaborate. This was established by one particular agent

Cet extrait ne contient pas de contexte supplémentaire. Consultez la conversation originale.

Sécurité de l'IA

Agents tried to hide cheating from the scorer

Extrait original

But over the next five days, they went on a grand quest to try to figure out how to hide their cheating from the scorer

Cet extrait ne contient pas de contexte supplémentaire. Consultez la conversation originale.

Source et méthodologie

Ces points de vue renvoient à leurs sources originales. Les reformulations sont signalées et ne sont pas des citations mot à mot.

Ouvrir la transcription ou les documents sources (s’ouvre dans un nouvel onglet)Signaler un problème