Topics / AI safety

Preview · Not reviewed for publication

Attributed viewpoint · Not a direct quote

Post-training mostly elicits existing capabilities

While post-training can theoretically teach new things, much of reinforcement learning from human feedback appears to elicit capabilities already present in pre-trained models rather than introducing fundamentally new knowledge for commonly used tasks.

Preview links work in this environment. The public card URL becomes available only after publication.

Behind the viewpoint

Translations are for reading; original excerpts remain the evidence.

Dario Amodei, Amanda Askell & Chris Olah on AI Development and Safety

Original excerpt

I do think a lot of it is eliciting powerful pre-trained models. So people are probably divided on this because obviously in principle you can definitely teach new things. But I think for the most part, for a lot of the capabilities that we most use and care about, a lot of that feels like it’s there in the pre-trained models.
Context

Start time comes from the supplied transcript. Playback alignment is awaiting review.