Topics / AI safety

Preview · Not reviewed for publication

Attributed viewpoint · Not a direct quote

Mechanistic interpretability reveals deception-related features

Researchers found multiple features related to deception, lying, withholding information, and power-seeking in neural networks. Forcing these features active causes the model to exhibit corresponding undesirable behaviors, though this area of research remains in early stages.

Preview links work in this environment. The public card URL becomes available only after publication.

Behind the viewpoint

Translations are for reading; original excerpts remain the evidence.

Dario Amodei, Amanda Askell & Chris Olah on AI Development and Safety

Original excerpt

we find quite a few features related to deception and lying. There’s one feature where it fires for people lying and being deceptive, and you force it active and Claude starts lying to you.
Context

Start time comes from the supplied transcript. Playback alignment is awaiting review.