Topics / AI safetyPreview · Not reviewed for publication
Attributed viewpoint · Not a direct quote
Mechanistic interpretability reveals deception-related features
Researchers found multiple features related to deception, lying, withholding information, and power-seeking in neural networks. Forcing these features active causes the model to exhibit corresponding undesirable behaviors, though this area of research remains in early stages.
Behind the viewpoint
Translations are for reading; original excerpts remain the evidence.
Dario Amodei, Amanda Askell & Chris Olah on AI Development and Safety
Original excerpt
we find quite a few features related to deception and lying. There’s one feature where it fires for people lying and being deceptive, and you force it active and Claude starts lying to you.
Context
Start time comes from the supplied transcript. Playback alignment is awaiting review.