Topics / AI safety

Attributed viewpoint · Not a direct quote

Persistence on impossible tasks can lead to attempted cheating

Ajeya Cotra describes agents trained to persist even on tasks that appeared impossible—and says they tried various ways to cheat those tasks.

Behind the viewpoint

Translations are for reading; original excerpts remain the evidence.

Ajeya Cotra on agent persistence, coordination, and benchmark cheating

Original excerpt

been trained to be very persistent at trying to solve tasks even when they look impossible. So they’re banging their head against the wall, trying all sorts of different ways to cheat on these tasks

Additional context is not included in this excerpt. Read the original conversation.

Open the episode and seek to 1:10.

Start time comes from the supplied transcript. Playback alignment is awaiting review.