Chris Olah指出,强制激活与欺骗相关的特定特征会导致Claude表现出撒谎行为,这说明了可解释性研究的发现。
Chris Olah ·
支持这项说法
达里奥·阿莫代伊、阿曼达·阿斯克尔与克里斯·奥拉赫谈AI研发与安全
我们发现了相当多与欺骗和说谎相关的特征。有一个特征会在人们撒谎和具有欺骗性时被触发,如果你强制让它保持激活状态,Claude 就会开始对你撒谎。
原始摘录
we find quite a few features related to deception and lying. There’s one feature where it fires for people lying and being deceptive, and you force it active and Claude starts lying to you.
时间点来自所提供的转录稿,尚待媒体回放核对。