Neural networks are not directly programmed but grown. Researchers design architectures as scaffolds and set loss objectives as directional targets, but the resulting system is more like a biological organism whose internal workings must be studied rather than code that was written.
Supporting evidence
Original excerpt
we don’t program, we don’t make them, we grow them. We have these neural network architectures that we design and we have these loss objectives that we create. And the neural network architecture, it’s kind of like a scaffold that the circuits grow on.
Mechanistic interpretability reveals deception-related features
Researchers found multiple features related to deception, lying, withholding information, and power-seeking in neural networks. Forcing these features active causes the model to exhibit corresponding undesirable behaviors, though this area of research remains in early stages.
Supporting evidence
Original excerpt
we find quite a few features related to deception and lying. There’s one feature where it fires for people lying and being deceptive, and you force it active and Claude starts lying to you.