Chris Olah notes that forcing specific deception-related features active causes Claude to exhibit lying behavior, illustrating interpretability findings.
Chris Olah ·
Supporting evidence
Dario Amodei, Amanda Askell & Chris Olah on AI Development and Safety
Original excerpt
we find quite a few features related to deception and lying. There’s one feature where it fires for people lying and being deceptive, and you force it active and Claude starts lying to you.
Start time comes from the supplied transcript. Playback alignment is awaiting review.
About this interpretation
The object-specific interpretation and Chinese translation were checked independently against the source. This is an AI semantic review, not playback verification. Reviewed Oct 5, 2026 · qwen3.8-max-0902
Report an issue