来源 / 访谈

原始对话

达里奥·阿莫代伊、阿曼达·阿斯克尔与克里斯·奥拉赫谈人工智能发展与安全

Lex Fridman Podcast · · 5:22:14

莱克斯·弗里德曼采访了Anthropic公司首席执行官达里奥·阿莫代伊,以及研究人员阿曼达·阿斯克尔和克里斯·奥拉赫。他们探讨了模型缩放动力学、数据限制、对齐策略、机制可解释性研究成果、强化学习效应、部署防护措施、监管设计考量,以及能力时间线预估。

Dario Amodei, Amanda Askell & Chris Olah on AI Development and Safety

翻译仅用于阅读理解;原始摘录仍为 source evidence。

一目了然

关键时刻9

精选观点,附带原文摘录及上下文。请在下方阅读完整文字稿。

AI基础设施

三种要素的线性缩放

原始摘录

almost like a chemical reaction, you have three ingredients in the chemical reaction and you need to linearly scale up the three ingredients. If you scale up one, not the others, you run out of the other reagents and the reaction stops.

翻译 · 非原文措辞

几乎就像一场化学反应,你有三种反应物,必须将这三种要素同步线性放大。如果你只放大其中一种,而不放大其余要素,就会耗尽其他反应物,反应便停止了。

上下文

AI安全

互联网数据质量约束

原始摘录

we simply run out of data. There’s only so much data on the internet, and there’s issues with the quality of the data. You can get hundreds of trillions of words on the internet, but a lot of it is repetitive or it’s search engine optimization drivel

翻译 · 非原文措辞

我们单纯会耗尽数据。互联网上的数据总量是有限的,且数据质量存在问题。你确实能从互联网上获取数百万亿词,但其中大量内容是重复的,或是搜索引擎优化类的废话。

上下文

AI安全

当前计算机使用模型需要边界与防护措施

原始摘录

It makes mistakes, it misclicks. We were careful to warn people, “Hey, you can’t just leave this thing to run on your computer for minutes and minutes. You got to give this thing boundaries and guardrails.” And I think that’s one of the reasons we released it first in an API form

翻译 · 非原文措辞

它会犯错,会误点击。我们谨慎提醒用户:‘嘿,你不能把它放在你的电脑上连续运行几分钟、几十分钟。你必须给它设定边界和防护措施。’我认为这正是我们首先以API形式发布它的原因之一。

上下文

AI安全

目标不当的人工智能监管可能引发对安全工作的反感情绪

原始摘录

if we get something in place that’s poorly targeted, that wastes a bunch of people’s time, what’s going to happen is people are going to say, “See, these safety risks, this is nonsense. I just had to hire 10 lawyers to fill out all these forms.

翻译 · 非原文措辞

如果我们推行一项目标不当的监管措施,白白浪费大量人员的时间,结果将是人们纷纷宣称:‘看吧,这些安全风险纯属无稽之谈。我刚被迫雇了10名律师填各种表格。’

上下文

决策制定

能力外推表明强大AI可能于2026–2027年到来

原始摘录

if you just eyeball the rate at which these capabilities are increasing, it does make you think that we’ll get there by 2026 or 2027. Again, lots of things could derail it. We could run out of data. We might not be able to scale clusters as much as we want.

翻译 · 非原文措辞

如果你仅凭肉眼观察这些能力提升的速度,的确会让你觉得我们将在2026年或2027年抵达这一阶段。再次强调,许多事情都可能使这一进程脱轨。我们可能耗尽数据;我们或许无法按预期规模扩展计算集群。

上下文

AI安全

角色塑造属于对齐工作,而不仅是产品考量

原始摘录

One thing I really like about the character work is from the outset it was seen as an alignment piece of work and not something like a product consideration

翻译 · 非原文措辞

我特别欣赏角色塑造工作的一点在于,它从一开始就被视为一项对齐工作,而非某种产品层面的考量。

上下文

AI安全

后训练主要激发已有能力

原始摘录

I do think a lot of it is eliciting powerful pre-trained models. So people are probably divided on this because obviously in principle you can definitely teach new things. But I think for the most part, for a lot of the capabilities that we most use and care about, a lot of that feels like it’s there in the pre-trained models.

翻译 · 非原文措辞

我确实认为其中很大一部分是在激发那些强大的预训练模型已具备的能力。因此,人们对此看法不一,因为原则上你当然可以教出全新内容;但我认为,就我们最常使用和最关心的诸多能力而言,其中大部分感觉早已存在于预训练模型之中。

上下文

AI基础设施

神经网络是‘生长’出来的,而非直接编程而成

原始摘录

we don’t program, we don’t make them, we grow them. We have these neural network architectures that we design and we have these loss objectives that we create. And the neural network architecture, it’s kind of like a scaffold that the circuits grow on.

翻译 · 非原文措辞

我们不是编程,也不是制造它们,而是培育它们。我们设计这些神经网络架构,也设定这些损失目标。而神经网络架构本身,某种程度上就像一个电路在其上生长的支架。

上下文

AI安全

机制可解释性揭示出与欺骗相关联的特征

原始摘录

we find quite a few features related to deception and lying. There’s one feature where it fires for people lying and being deceptive, and you force it active and Claude starts lying to you.

翻译 · 非原文措辞

我们发现了相当多与欺骗和说谎相关的特征。其中有一个特征会在人们撒谎和具有欺骗性时被触发;当你强制激活它时,Claude就开始对你撒谎。

上下文

收听采访录音

原始出版方音频。文字稿的时间戳可能基于不同版本的视频。

参与者

Lex Fridman主持人Dario Amodei嘉宾Amanda Askell嘉宾Chris Olah嘉宾

章节导航

出版方时间戳;播放对齐功能正在开发中。

  1. 0:00引言
  2. 3:14缩放定律
  3. 12:20大语言模型缩放的限制
  4. 20:46与OpenAI、Google、xAI、Meta的竞争
  5. 26:08Claude
  6. 29:44Opus 3.5
  7. 34:30Sonnet 3.5
  8. 37:49Claude 4.0
  9. 42:02对Claude的批评
  10. 54:49人工智能安全等级
  11. 1:05:37ASL-3与ASL-4
  12. 1:09:40计算机使用
  13. 1:19:36政府对人工智能的监管
  14. 1:38:25组建优秀团队
  15. 1:47:14后训练
  16. 1:52:39宪法式人工智能(Constitutional AI)
  17. 1:58:06充满慈爱的机器(Machines of Loving Grace)
  18. 2:17:11通用人工智能(AGI)时间线
  19. 2:29:46编程
  20. 2:36:45生命的意义
  21. 2:42:44阿曼达·阿斯克尔
  22. 2:45:21面向非技术人员的编程建议
  23. 2:49:10与Claude对话
  24. 3:05:38提示工程(Prompt engineering)
  25. 3:14:15后训练
  26. 3:18:52宪法式人工智能(Constitutional AI)
  27. 3:23:48系统提示(System prompts)
  28. 3:29:55Claude是否正在变笨?
  29. 3:41:56角色训练
  30. 3:42:49真理的本质
  31. 3:47:32最优失败率
  32. 3:54:43人工智能意识
  33. 4:09:15通用人工智能(AGI)
  34. 4:17:45克里斯·奥拉赫
  35. 4:22:40特征、电路与普适性
  36. 4:40:17叠加态(Superposition)
  37. 4:51:16单义性(Monosemanticity)
  38. 4:57:58单义性的缩放
  39. 5:06:56神经网络的宏观行为
  40. 5:11:51神经网络之美

完整文字记录

AI 翻译为中文,英文原文将一并保留。

AI 已审阅;尚待人工独立核查及回放验证。

Dario Amodei0:00

如果你外推我们迄今所见的那些曲线,对吧?假如你说:‘嗯,我不知道,我们目前正迈向博士水平,而去年还停留在本科生水平,前年则仅相当于高中生水平。’当然,你尽可就具体任务及评判标准展开争论。‘我们仍缺少某些模态,但这些模态正被陆续加入’——比如计算机使用能力已被加入,图像生成能力也已被加入。倘若你粗略估算这些能力提升的速度,的确会让人觉得我们有望在2026年或2027年达成目标。

英文原文

If you extrapolate the curves that we’ve had so far, right? If you say, “Well, I don’t know, we’re starting to get to PhD level, and last year we were at undergraduate level, and the year before we were at the level of a high school student,” again, you can quibble with what tasks and for what. “We’re still missing modalities, but those are being added,” like computer use was added, like image generation has been added. If you just kind of eyeball the rate at which these capabilities are increasing, it does make you think that we’ll get there by 2026 or 2027.

Dario Amodei0:31

我认为,仍存在一些世界,在其中此事一百年内都不会发生。但这类世界的数量正迅速减少。我们正快速耗尽真正令人信服的阻碍因素,真正有力的理由来说明此事在未来几年内不会发生。模型规模扩张的速度极快。我们今天就能做到:先构建一个模型,再部署成千上万、甚至数以万计的实例。我认为,至多两到三年内,无论这些超强人工智能是否已问世,计算集群的规模都将发展到足以支持部署数百万个此类模型的程度。

英文原文

I think there are still worlds where it doesn’t happen in 100 years. The number of those worlds is rapidly decreasing. We are rapidly running out of truly convincing blockers, truly compelling reasons why this will not happen in the next few years. The scale-up is very quick. We do this today, we make a model, and then we deploy thousands, maybe tens of thousands of instances of it. I think by the time, certainly within two to three years, whether we have these super powerful AIs or not, clusters are going to get to the size where you’ll be able to deploy millions of these.

Dario Amodei1:03

我对‘意义’持乐观态度。我担忧的是经济影响以及权力的集中。事实上,我更担忧的正是权力的滥用。

英文原文

I am optimistic about meaning. I worry about economics and the concentration of power. That’s actually what I worry about more, the abuse of power.

Lex Fridman1:14

而人工智能增加了世界上的权力总量。倘若将这种权力加以集中并滥用,便可能造成难以估量的损害。

英文原文

And AI increases the amount of power in the world. And if you concentrate that power and abuse that power, it can do immeasurable damage.

Dario Amodei1:22

是的,这非常可怕。非常可怕。

英文原文

Yes, it’s very frightening. It’s very frightening.

Lex Fridman1:27

以下是一段与Anthropic公司首席执行官达里奥·阿莫代伊(Dario Amodei)的对话。Anthropic公司开发了Claude模型,该模型当前常居多数大语言模型(LLM)基准测试排行榜榜首。此外,达里奥及Anthropic团队一直公开倡导严肃对待人工智能安全议题,并持续就该主题及其他相关领域发表大量引人入胜的研究成果。

英文原文

The following is a conversation with Dario Amodei, CEO of Anthropic, the company that created Claude, that is currently and often at the top of most LLM benchmark leader boards. On top of that, Dario and the Anthropic team have been outspoken advocates for taking the topic of AI safety very seriously. And they have continued to publish a lot of fascinating AI on this and other topics.

Lex Fridman1:55

随后,我还邀请了Anthropic另外两位杰出人士参与对话。首先是阿曼达·阿斯克尔(Amanda Askell),她是一名研究员,专注于Claude模型的对齐(alignment)与微调(fine-tuning),包括Claude角色与人格的设计。有几位同事告诉我,她与Claude对话的次数可能超过Anthropic内部任何一位人类。因此,她无疑是探讨提示工程(prompt engineering)以及如何充分发挥Claude潜力之实用建议的绝佳人选。

英文原文

I’m also joined afterwards by two other brilliant people from Anthropic. First Amanda Askell, who is a researcher working on alignment and fine-tuning of Claude, including the design of Claude’s character and personality. A few folks told me she has probably talked with Claude more than any human at Anthropic. So she was definitely a fascinating person to talk to about prompt engineering and practical advice on how to get the best out of Claude.

Lex Fridman2:27

之后,克里斯·奥拉赫(Chris Olah)也前来参与了交流。他是“机制可解释性”(mechanistic interpretability)这一领域的先驱之一。该领域是一系列激动人心的研究工作,旨在对神经网络进行逆向工程,以探明其内部运作机制,即通过分析网络内部神经元激活模式来推断其行为。这是一种极具前景的方法,可用于保障未来超级智能人工智能系统的安全性。例如,可通过监测神经元激活状态,检测模型是否正试图欺骗与其对话的人类。

英文原文

After that, Chris Olah stopped by for a chat. He’s one of the pioneers of the field of mechanistic interpretability, which is an exciting set of efforts that aims to reverse engineering neural networks, to figure out what’s going on inside, inferring behaviors from neural activation patterns inside the network. This is a very promising approach for keeping future super-intelligent AI systems safe. For example, by detecting from the activations when the model is trying to deceive the human it is talking to.

Lex Fridman3:03

这是Lex Fridman播客节目。如欲支持本节目,请查看简介中的赞助商信息。现在,亲爱的朋友们,让我们欢迎达里奥·阿莫代伊。

英文原文

This is the Lex Fridman podcast. To support it, please check out our sponsors in the description. And now, dear friends, here’s Dario Amodei.

Lex Fridman3:14

我们先从‘缩放定律’(scaling laws)与‘缩放假说’(scaling hypothesis)这一宏大构想谈起。它是什么?其历史渊源如何?我们当下又处于什么位置?

英文原文

Let’s start with a big idea of scaling laws and the scaling hypothesis. What is it? What is its history, and where do we stand today?

Dario Amodei3:22

我只能依据自身经历来描述它。我在人工智能领域已有约十年从业经验,而这一现象在我入行之初便已引起我的注意。我最早于2014年底加入人工智能领域,在百度与吴恩达(Andrew Ng)共事。我们最初着手的工作是语音识别系统。彼时,深度学习尚属新兴技术,虽已取得诸多进展,但业内普遍认为:‘我们尚未掌握成功所需的算法。我们仅能匹配极小一部分能力;我们在算法层面尚需发现大量新知;我们仍未找到模拟人类大脑运作方式的图景。’

英文原文

So I can only describe it as it relates to my own experience, but I’ve been in the AI field for about 10 years and it was something I noticed very early on. So I first joined the AI world when I was working at Baidu with Andrew Ng in late 2014, which is almost exactly 10 years ago now. And the first thing we worked on, was speech recognition systems. And in those days I think deep learning was a new thing. It had made lots of progress, but everyone was always saying, “We don’t have the algorithms we need to succeed. We are only matching a tiny fraction. There’s so much we need to discover algorithmically. We haven’t found the picture of how to match the human brain.”

Dario Amodei4:05

某种程度上,这或许是一种幸运,近乎初学者的好运。我当时刚踏入该领域,是个新人。我审视了我们用于语音识别的神经网络——循环神经网络(RNN),心想:‘如果让它们更大些、增加更多层,同时随模型规模扩大一并扩充训练数据,结果会怎样?’当时我仅将这些视为可独立调节的旋钮。我注意到,随着数据量增加、模型规模扩大、训练时间延长,模型性能持续提升。虽然当时我并未精确量化这些因素,但与同事们共同形成的非正式共识是:投入的数据越多、算力越强、训练越充分,模型表现就越好。

英文原文

And in some ways it was fortunate, you can have almost beginner’s luck. I was like a newcomer to the field. And I looked at the neural net that we were using for speech, the recurrent neural networks, and I said, “I don’t know, what if you make them bigger and give them more layers? And what if you scale up the data along with this?” I just saw these as independent dials that you could turn. And I noticed that the models started to do better and better as you gave them more data, as you made the models larger, as you trained them for longer. And I didn’t measure things precisely in those days, but along with colleagues, we very much got the informal sense that the more data and the more compute and the more training you put into these models, the better they perform.

Dario Amodei4:51

起初,我的想法是:‘嘿,也许这只适用于语音识别系统,或许只是某一特定领域特有的偶然现象。’直到2017年,当我首次看到GPT-1的实验结果时,我才恍然领悟:语言处理很可能是我们能够实现这一目标的领域。我们可以获取万亿级词汇量的语言数据,并以此进行训练。而当时我们训练的模型规模极小,仅需一至八块GPU即可完成;如今,我们已在数以万计的GPU上开展训练任务,且很快将扩展至数十万块GPU。

英文原文

And so initially my thinking was, “Hey, maybe that is just true for speech recognition systems. Maybe that’s just one particular quirk, one particular area.” I think it wasn’t until 2017 when I first saw the results from GPT-1 that it clicked for me that language is probably the area in which we can do this. We can get trillions of words of language data, we can train on them. And the models we were trained in those days were tiny. You could train them on one to eight GPUs, whereas now we train jobs on tens of thousands, soon going to hundreds of thousands of GPUs.

Dario Amodei5:28

因此,当我将这两点联系起来时,才真正豁然开朗。像你此前采访过的伊利亚·苏茨克维(Ilya Sutskever)等少数人,也持有类似观点。他或许是首位提出该观点者,尽管我认为当时亦有若干人几乎同步得出了相似结论,对吧?理查德·萨顿(Rich Sutton)曾提出‘苦涩的教训’(bitter lesson),格温(Gwern)也曾撰文论述缩放假说。但我认为,真正让我确信无疑的时间点,大致介于2014至2017年间——那时我真正坚定了信念:‘嘿,只要持续扩大模型规模,我们就将有能力应对极其宽泛的认知任务。’

英文原文

And so when I saw those two things together, and there were a few people like Ilya Sudskever who you’ve interviewed, who had somewhat similar views. He might’ve been the first one, although I think a few people came to similar views around the same time, right? There was Rich Sutton’s bitter lesson, Gwern wrote about the scaling hypothesis. But I think somewhere between 2014 and 2017 was when it really clicked for me, when I really got conviction that, “Hey, we’re going to be able to these incredibly wide cognitive tasks if we just scale up the models.”

Dario Amodei6:03

而在每一次缩放阶段,总会出现各种质疑声音。当我初次听到这些质疑时,坦率地说,我甚至以为:‘大概错的是我,而领域内所有专家才是对的。他们比我更了解实际情况,对吧?’例如乔姆斯基(Chomsky)提出的论点:‘你或许能掌握句法,却无法获得语义。’还有这样一种观点:‘你可以让单个句子通顺合理,却无法让整段文字逻辑自洽。’而我们今天最新的质疑则是:‘我们将面临数据枯竭,或数据质量不足,抑或模型本身不具备推理能力。’

英文原文

And at every stage of scaling, there are always arguments. And when I first heard them honestly, I thought, “Probably I’m the one who’s wrong and all these experts in the field are right. They know the situation better than I do, right?” There’s the Chomsky argument about, “You can get syntactics but you can’t get semantics.” There was this idea, “Oh, you can make a sentence make sense, but you can’t make a paragraph make sense.” The latest one we have today is, “We’re going to run out of data, or the data isn’t high quality enough or models can’t reason.”

Dario Amodei6:34

而每一次,我们总能设法绕过这些障碍,或者,缩放本身恰恰就是那条出路。有时是前者,有时是后者。因此,我如今虽仍认为一切始终充满不确定性——毕竟我们唯一能依赖的,不过是归纳推理,用过去十年的经验去推测未来两年的发展趋势——但这段‘剧情’我已看过太多遍,类似的故事已反复上演多次,使我深信缩放进程很可能将持续下去,且其中蕴含某种我们尚未从理论层面真正阐明的‘魔力’。

英文原文

And each time, every time, we manage to either find a way around or scaling just is the way around. Sometimes it’s one, sometimes it’s the other. And so I’m now at this point, I still think it’s always quite uncertain. We have nothing but inductive inference to tell us that the next two years are going to be like the last 10 years. But I’ve seen the movie enough times, I’ve seen the story happen for enough times to really believe that probably the scaling is going to continue, and that there’s some magic to it that we haven’t really explained on a theoretical basis yet.

Lex Fridman7:10

当然,此处所说的‘缩放’,指的是更大的网络、更多的数据、更强的算力?

英文原文

And of course the scaling here is bigger networks, bigger data, bigger compute?

Dario Amodei7:16

是的。

英文原文

Yes.

Lex Fridman7:17

三者皆是?

英文原文

All of those?

Dario Amodei7:17

尤其是指网络规模、训练时长及数据量的线性增长。因此,这三者几乎如同化学反应中的三种原料,必须同步线性放大。若仅放大其中一项,其余两项未同步放大,则会因其他‘反应物’耗尽而导致‘反应’中止;但若三者同步放大,则‘反应’方可持续推进。

英文原文

In particular, linear scaling up of bigger networks, bigger training times and more and more data. So all of these things, almost like a chemical reaction, you have three ingredients in the chemical reaction and you need to linearly scale up the three ingredients. If you scale up one, not the others, you run out of the other reagents and the reaction stops. But if you scale up everything in series, then the reaction can proceed.

Lex Fridman7:45

当然,如今既然已形成这样一门经验性的科学/艺术,我们便可将其应用于其他更精细的领域,例如将缩放定律应用于可解释性研究,或应用于后训练(post-training)阶段;又或单纯观察某项指标究竟如何随规模变化。但所谓核心缩放定律,或者说其底层的缩放假说,本质上是否指向‘更大的网络、更多的数据将导向更高智能’?

英文原文

And of course now that you have this kind of empirical science/art, you can apply it to other more nuanced things like scaling laws applied to interpretability or scaling laws applied to post-training. Or just seeing how does this thing scale. But the big scaling law, I guess the underlying scaling hypothesis has to do with big networks, big data leads to intelligence?

Dario Amodei8:09

是的,我们已在语言之外的众多领域验证了缩放定律。最初展示该规律的论文发布于2020年初,首次在语言领域证实了这一点。随后在2020年末,又有研究在图像、视频、文本生成图像、图像生成文本、数学等领域同样揭示了相同规律。您说得对,如今又出现了后训练等新阶段,以及新型推理模型。而在我们已测量的所有此类案例中,均观察到了类似的缩放定律。

英文原文

Yeah, we’ve documented scaling laws in lots of domains other than language. So initially the paper we did that first showed it, was in early 2020, where we first showed it for language. There was then some work late in 2020 where we showed the same thing for other modalities like images, video, text to image, image to text, math. They all had the same pattern. And you’re right, now there are other stages like post-training or there are new types of reasoning models. And in all of those cases that we’ve measured, we see similar types of scaling laws.

Lex Fridman8:48

稍带一点哲学意味的问题:关于为何网络规模与数据规模越大越好,您的直觉判断是什么?为何这会导致模型更趋智能?

英文原文

A bit of a philosophical question, but what’s your intuition about why bigger is better in terms of network size and data size? Why does it lead to more intelligent models?

Dario Amodei9:00

所以在我之前作为生物物理学家的职业生涯中……我本科读的是物理学,然后在研究生阶段转向了生物物理学。因此,我回溯自己作为物理学家所掌握的知识——实际上,这方面的专业知识远不如我在Anthropic的一些同事深厚。这里有个概念叫作‘1/f噪声’和‘1/x分布’:通常,就像当你叠加大量自然过程时会得到高斯分布一样,当你叠加大量具有不同分布特性的自然过程时……如果你拿一个探针连接到一个电阻上,该电阻中热噪声的分布就随频率呈1/f关系。这是一种某种天然的收敛分布。

英文原文

So in my previous career as a biophysicist… So I did a physics undergrad and then biophysics in grad school. So I think back to what I know as a physicist, which is actually much less than what some of my colleagues at Anthropic have in terms of expertise in physics. There’s this concept called the one over F noise and one over X distributions, where often, just like if you add up a bunch of natural processes, you get a Gaussian, if you add up a bunch of differently-distributed natural processes… If you take a probe and hook it up to a resistor, the distribution of the thermal noise in the resistor goes as one over the frequency. It’s some kind of natural convergent distribution.

Dario Amodei9:50

而我认为其本质在于:如果你观察大量由具备多尺度特征的自然过程所产生的现象,它们并非高斯分布(高斯分布属于相对窄幅集中的分布);但若我观察导致电噪声的大尺度与小尺度涨落,它们便呈现出这种衰减型的1/x分布。因此,当我思考物理世界或语言中的模式时,如果考虑语言中的模式,就会发现一些极其简单的模式:某些词远比其他词更常见,比如‘the’;接着是基本的名词—动词结构;再接着是名词与动词需保持一致、需彼此协调这一事实;然后是更高层级的句子结构;再往上则是段落的主题结构。因此,这种逐层递进的结构意味着:你可以设想,随着网络规模增大,它首先捕获的是那些最简单相关性、最简单模式,而其余模式则构成一条长长的尾部。

英文原文

And I think what it amounts to, is that if you look at a lot of things that are produced by some natural process that has a lot of different scales, not a Gaussian, which is kind of narrowly distributed, but if I look at large and small fluctuations that lead to electrical noise, they have this decaying one over X distribution. And so now I think of patterns in the physical world or in language. If I think about the patterns in language, there are some really simple patterns, some words are much more common than others, like the. Then there’s basic noun-verb structure. Then there’s the fact that nouns and verbs have to agree, they have to coordinate. And there’s the higher-level sentence structure. Then there’s the thematic structure of paragraphs. And so the fact that there’s this regressing structure, you can imagine that as you make the networks larger, first they capture the really simple correlations, the really simple patterns, and there’s this long tail of other patterns.

Dario Amodei10:49

而如果这条长尾中的其他模式,正如电阻等物理过程中出现的1/f噪声那样,确实非常平滑,那么你就可以设想:随着网络规模增大,它逐步捕捉到该分布中越来越多的内容。而这种平滑性,便会反映在模型的预测能力与整体性能表现上。

英文原文

And if that long tail of other patterns is really smooth like it is with the one over F noise in physical processes like resistors, then you can imagine as you make the network larger, it’s kind of capturing more and more of that distribution. And so that smoothness gets reflected in how well the models are at predicting and how well they perform.

Dario Amodei11:10

语言是一种演化而来的过程。我们发展出了语言,拥有常用词与不常用词,拥有常用表达与不常用表达,拥有被频繁使用的观念与陈词滥调,也拥有新颖的想法。这一过程历经数百万年,在人类演化过程中逐步形成。因此,一种猜测——纯属推测——是:这些观念的分布本身可能就呈现某种长尾分布。

英文原文

Language is an evolved process. We’ve developed language, we have common words and less common words. We have common expressions and less common expressions. We have ideas, cliches, that are expressed frequently, and we have novel ideas. And that process has developed, has evolved with humans over millions of years. And so the guess, and this is pure speculation, would be that there’s some kind of long tail distribution of the distribution of these ideas.

Lex Fridman11:41

所以,这里既有长尾,也有你正在构建的概念层级的高度。因此,网络越大,理论上你就拥有越高的……

英文原文

So there’s the long tail, but also there’s the height of the hierarchy of concepts that you’re building up. So the bigger the network, presumably you have a higher capacity to-

Dario Amodei11:50

没错。如果你用一个小型网络,你只能学到那些常见内容。如果我用一个极小的神经网络,它很擅长理解一个句子必须包含动词、形容词、名词,但在判断这些动词、形容词、名词具体应为何物,以及它们是否合理搭配方面却表现极差。如果我仅将网络稍加扩大,它便能很好地处理这类问题;接着它突然又擅长处理句子了,却仍无法胜任段落层面的任务。因此,这些更稀有、更复杂的模式,会随着我为网络增加更多容量而被逐步捕获。

英文原文

Exactly. If you have a small network, you only get the common stuff. If I take a tiny neural network, it’s very good at understanding that a sentence has to have verb, adjective, noun, but it’s terrible at deciding what those verb adjective and noun should be and whether they should make sense. If I make it just a little bigger, it gets good at that, then suddenly it’s good at the sentences, but it’s not good at the paragraphs. And so these rarer and more complex patterns get picked up as I add more capacity to the network.

Lex Fridman12:20

那么,一个自然而然的问题便是:这种增长是否存在上限?

英文原文

Well, the natural question then is what’s the ceiling of this?

Dario Amodei12:24

是的。

英文原文

Yeah.

Lex Fridman12:24

现实世界究竟有多复杂?其中可供学习的内容到底有多少?

英文原文

How complicated and complex is the real world? How much is the stuff is there to learn?

Dario Amodei12:30

我想,我们当中没有任何人知道这个问题的答案。我强烈的直觉是:在人类能力水平之下,并不存在这样的上限。人类自身能够理解各种各样的模式,因此这让我相信:如果我们持续扩大这些模型的规模,并同步发展出新的训练方法与规模化方法,至少能达到人类目前所达到的水平。接下来的问题是:人类之上,是否还存在进一步理解的空间?人工智能是否有可能比人类更聪明、更具洞察力?我猜测,答案必然取决于具体领域。

英文原文

I don’t think any of us knows the answer to that question. My strong instinct would be that there’s no ceiling below the level of humans. We humans are able to understand these various patterns. And so that makes me think that if we continue to scale up these models to kind of develop new methods for training them and scaling them up, that will at least get to the level that we’ve gotten to with humans. There’s then a question of how much more is it possible to understand than humans do? How much is it possible to be smarter and more perceptive than humans? I would guess the answer has got to be domain-dependent.

Dario Amodei13:09

如果我聚焦于生物学这类领域,并参考我写过的那篇题为《仁慈之机》(Machines of Loving Grace)的文章,我的感觉是:人类正艰难地试图理解生物学的复杂性。如果你去斯坦福大学、哈佛大学或加州大学伯克利分校,你会发现整座院系都在研究免疫系统或代谢通路,而每位研究人员仅能理解其中极小一部分,仅专精于某个细分方向;他们还在努力将自己的知识与其他人的知识整合起来。因此,我有一种直觉:在顶端,人工智能仍有巨大的提升空间。

英文原文

If I look at an area like biology, and I wrote this essay, Machines of Loving Grace, it seems to me that humans are struggling to understand the complexity of biology. If you go to Stanford or to Harvard or to Berkeley, you have whole departments of folks trying to study the immune system or metabolic pathways, and each person understands only a tiny bit, a part of it, specializes. And they’re struggling to combine their knowledge with that of other humans. And so I have an instinct that there’s a lot of room at the top for AIs to get smarter.

Dario Amodei13:46

如果我想到物理世界中的材料,或者人类之间的冲突调解之类的问题,我的意思是,其中某些问题或许并非不可解,只是难度极大;也可能在某些事情上,人类所能达到的水平终究存在极限。就像语音识别一样,我所能听清你讲话的清晰度终究有限。因此,我认为在某些领域,天花板可能非常接近人类已达到的水平;而在另一些领域,天花板则可能遥不可及。我想,我们只有真正构建出这些系统之后,才能知晓答案。事先很难准确预判,我们只能推测,却无法确信。

英文原文

If I think of something like materials in the physical world, or addressing conflicts between humans or something like that, I mean it may be there’s only some of these problems are not intractable, but much harder. And it may be that there’s only so well you can do at some of these things. Just like with speech recognition, there’s only so clear I can hear your speech. So I think in some areas there may be ceilings that are very close to what humans have done. In other areas, those ceilings may be very far away. I think we’ll only find out when we build these systems. It’s very hard to know in advance. We can speculate, but we can’t be sure.

Lex Fridman14:26

而在某些领域,天花板或许与人类官僚体系之类因素有关,正如你在文中所论述的那样。

英文原文

And in some domains, the ceiling might have to do with human bureaucracies and things like this, as you write about.

Dario Amodei14:31

是的。

英文原文

Yes.

Lex Fridman14:31

因此,人类从根本上必须参与其中,这才是造成天花板的原因,而非智能本身的局限。

英文原文

So humans fundamentally has to be part of the loop. That’s the cause of the ceiling, not maybe the limits of the intelligence.

Dario Amodei14:38

是的,我认为在许多情况下,从理论上讲,技术本可发展得极快。例如,我们在生物学领域可能发明的一切成果,但请记住,我们必须通过临床试验体系,才能真正将这些成果应用于人类。我认为这一体系混杂着两类因素:一类是官僚体制中不必要的部分,另一类则是旨在维护社会完整性的部分。整个挑战恰恰在于:我们难以分辨现状究竟如何,难以区分哪些是前者、哪些是后者。

英文原文

Yeah, I think in many cases, in theory, technology could change very fast. For example, all the things that we might invent with respect to biology, but remember, there’s a clinical trial system that we have to go through to actually administer these things to humans. I think that’s a mixture of things that are unnecessary in bureaucratic and things that kind of protect the integrity of society. And the whole challenge is that it’s hard to tell what’s going on. It’s hard to tell which is which.

Dario Amodei15:11

就药物研发而言,我的看法是:我们当前进展太慢,也过于保守。但当然,倘若在这些事情上犯错,过于鲁莽的确可能危及人们的生命。因此,至少某些人类制度确实在保护民众。关键在于找到平衡点。我强烈怀疑,这种平衡总体上更偏向于希望加快进程,但平衡本身确实存在。

英文原文

I think in terms of drug development, my view is that we’re too slow and we’re too conservative. But certainly if you get these things wrong, it’s possible to risk people’s lives by being too reckless. And so at least some of these human institutions are in fact protecting people. So it’s all about finding the balance. I strongly suspect that balance is kind of more on the side of wishing to make things happen faster, but there is a balance.

Lex Fridman15:39

如果我们真的触及极限,如果缩放定律真的开始放缓,你认为原因会是什么?是算力受限?数据受限?还是别的原因?抑或是创意受限?

英文原文

If we do hit a limit, if we do hit a slowdown in the scaling laws, what do you think would be the reason? Is it compute-limited, data-limited? Is it something else? Idea limited?

Dario Amodei15:51

有几点需要说明:我们现在讨论的是在模型尚未达到人类水平与人类技能之前就触及极限的情形。我认为当前一个广受关注、且确实可能成为我们遭遇的瓶颈的因素是:我们单纯耗尽了数据。互联网上的数据总量终归有限,而且数据质量也存在问题。你或许能从互联网上获取数百万万亿字的文本,但其中大量内容重复冗余,或充斥着搜索引擎优化(SEO)垃圾信息;甚至未来,其中部分内容可能干脆就是由人工智能自身生成的文本。因此,我认为这种方式所能产出的数据存在固有上限。

英文原文

So a few things, now we’re talking about hitting the limit before we get to the level of humans and the skill of humans. So I think one that’s popular today, and I think could be a limit that we run into, like most of the limits, I would bet against it, but it’s definitely possible, is we simply run out of data. There’s only so much data on the internet, and there’s issues with the quality of the data. You can get hundreds of trillions of words on the internet, but a lot of it is repetitive or it’s search engine optimization drivel, or maybe in the future it’ll even be text generated by AIs itself. And so I think there are limits to what can be produced in this way.

Dario Amodei16:34

话虽如此,我们——我估计其他公司也是如此——正在探索生成合成数据的方法:即利用现有模型生成更多同类数据,甚至凭空生成全新数据。以DeepMind的AlphaGo Zero为例,他们让一个完全不会下围棋的程序,仅通过自我对弈,便一路成长为超越人类水平的棋手;AlphaGo Zero版本中根本无需任何人类对局样本数据。

英文原文

That said, we, and I would guess other companies, are working on ways to make data synthetic, where you can use the model to generate more data of the type that you have already, or even generate data from scratch. If you think about what was done with DeepMind’s AlphaGo Zero, they managed to get a bot all the way from no ability to play Go whatsoever to above human level, just by playing against itself. There was no example data from humans required in the AlphaGo Zero version of it.

Dario Amodei17:07

另一个方向则是推理模型,它们采用思维链(chain-of-thought)方式,会在推理过程中暂停、反思自身思路。这种方式本质上是将合成数据与强化学习相结合。因此,我猜测:借助上述任一方法,我们都能绕过数据限制;或者,也可能存在其他尚未发掘的数据来源。我们还可以观察到:即便数据本身毫无问题,随着模型不断放大,其性能提升也可能戛然而止。此前人们曾可靠地观察到模型性能持续提升,但这种提升也可能在某个未知原因的作用下,于某一点彻底停止。

英文原文

The other direction of course, is these reasoning models that do chain of thought and stop to think and reflect on their own thinking. In a way that’s another kind of synthetic data coupled with reinforcement learning. So my guess is with one of those methods, we’ll get around the data limitation or there may be other sources of data that are available. We could just observe that, even if there’s no problem with data, as we start to scale models up, they just stopped getting better. It seemed to be a reliable observation that they’ve gotten better, that could just stop at some point for a reason we don’t understand.

Dario Amodei17:43

答案或许是:我们需要发明某种全新的架构。过去曾出现过一些问题,例如模型的数值稳定性问题——当时看起来性能似乎已趋于平稳,但事实上,当我们找到正确的‘解封器’后,性能并未真正停滞。因此,或许我们需要某种新的优化方法或某种新技术来突破瓶颈。截至目前,我尚未看到任何相关证据;但如果进展真的放缓,这或许会成为一个原因。

英文原文

The answer could be that we need to invent some new architecture. There have been problems in the past with say, numerical stability of models where it looked like things were leveling off, but actually when we found the right unblocker, they didn’t end up doing so. So perhaps there’s some new optimization method or some new technique we need to unblock things. I’ve seen no evidence of that so far, but if things were to slow down, that perhaps could be one reason.

Lex Fridman18:15

那么,算力是否存在极限?也就是说,建造规模越来越大、成本越来越高的数据中心,这种做法本身是否不可持续?

英文原文

What about the limits of compute, meaning the expensive nature of building bigger and bigger data centers?

Dario Amodei18:23

目前,我认为大多数前沿模型公司——我估计——大致处于约10亿参数量级,上下浮动三倍左右。这些就是当前已存在或正在训练中的模型。我认为明年我们将迈入数十亿参数量级,而到2026年,可能将突破100亿参数量级。到2027年,业界雄心勃勃地计划建设价值千亿美元级别的算力集群。我认为这一切确实会发生。各方都展现出极强的决心,要在国内建成这些算力设施,我估计它们最终真会落地。

英文原文

So right now, I think most of the frontier model companies, I would guess, are operating in roughly 1 billion scale, plus or minus a factor of three. Those are the models that exist now or are being trained now. I think next year we’re going to go to a few billion, and then 2026, we may go to above 10 billion. And probably by 2027, their ambitions to build hundred billion dollar clusters. And I think all of that actually will happen. There’s a lot of determination to build the compute, to do it within this country, and I would guess that it actually does happen.

Dario Amodei19:02

现在,如果我们真达到千亿美元级别的算力投入,那仍然不够,规模依然不足;此时,我们要么需要进一步扩大规模,要么必须开发出更高效的方法,从而改变这条增长曲线。在所有这些因素中,我之所以对强大AI如此快速到来持乐观态度,原因之一就在于:仅凭对曲线未来几个点的外推,我们便正迅速逼近人类水平的能力。

英文原文

Now, if we get to a hundred billion, that’s still not enough compute, that’s still not enough scale, then either we need even more scale, or we need to develop some way of doing it more efficiently of shifting the curve. I think between all of these, one of the reasons I’m bullish about powerful AI happening so fast, is just that if you extrapolate the next few points on the curve, we’re very quickly getting towards human level ability.

Dario Amodei19:28

我们开发的一些新模型,以及来自其他公司的某些推理模型,已开始达到我所称的‘博士’或‘专业级’水平。以编程能力为例:我们最新发布的Sonnet 3.5(即新版或更新版)在SWE-bench基准测试中得分约为50%。SWE-bench是一组涵盖大量真实世界专业软件工程任务的评测。今年年初,该基准的最高水平大约仅为3%或4%。也就是说,短短10个月内,我们在这一任务上的表现就从3%跃升至50%。而我认为再过一年,得分很可能达到90%——当然,我也不能确定,甚至可能用不了那么久。

英文原文

Some of the new models that we developed, some reasoning models that have come from other companies, they’re starting to get to what I would call the PhD or professional level. If you look at their coding ability, the latest model we released, Sonnet 3.5, the new or updated version, it gets something like 50% on SWE-bench. And SWE-bench is an example of a bunch of professional real-world software engineering tasks. At the beginning of the year, I think the state of the art was 3 or 4%. So in 10 months we’ve gone from 3% to 50% on this task. And I think in another year we’ll probably be at 90%. I mean, I don’t know, but might even be less than that.

Dario Amodei20:11

我们在研究生级别的数学、物理和生物学等领域也观察到了类似现象,例如OpenAI的o1模型。因此,若仅按我们目前已掌握的技能水平继续外推,我认为,只要沿这条直线持续外推几年,这些模型在各项能力上就将超越人类所能达到的最高专业水准。不过,这条外推曲线是否会持续下去?你已指出、我也指出过许多可能导致其无法延续的原因。但倘若该外推曲线的确延续下去,那这就是我们当前所处的轨迹。

英文原文

We’ve seen similar things in graduate-level math, physics, and biology from models like OpenAi’s o1. So if we just continue to extrapolate this in terms of skill that we have, I think if we extrapolate the straight curve, within a few years, we will get to these models being above the highest professional level in terms of humans. Now, will that curve continue? You’ve pointed to, and I’ve pointed to a lot of possible reasons why that might not happen. But if the extrapolation curve continues, that is the trajectory we’re on.

Lex Fridman20:46

Anthropic拥有若干竞争对手。您能否谈谈您对整个竞争格局的看法?比如OpenAI、谷歌、XAI、Meta等公司。广义而言,要在这一领域‘胜出’,究竟需要什么?

英文原文

So Anthropic has several competitors. It’d be interesting to get your sort of view of it all. OpenAI, Google, XAI, Meta. What does it take to win in the broad sense of win in this space?

Dario Amodei20:58

是的,我想先厘清几点。Anthropic的使命,本质上是努力确保这一切顺利发展。我们提出了一种名为‘竞相向善’(Race to the Top)的变革理论。‘竞相向善’旨在通过树立榜样,推动其他参与者也采取正确行动。它并非仅仅关乎成为‘好人’,而是要构建一种机制,使我们所有人皆可成为‘好人’。

英文原文

Yeah, so I want to separate out a couple things, right? Anthropic’s mission is to kind of try to make this all go well. And we have a theory of change called Race to the Top. Race to the Top is about trying to push the other players to do the right thing by setting an example. It’s not about being the good guy, it’s about setting things up so that all of us can be the good guy.

Dario Amodei21:24

我举几个例子说明。Anthropic创立早期,我们的联合创始人之一克里斯·奥拉赫(Chris Olah)——我相信您很快就会采访他——他是‘机制可解释性’(mechanistic interpretability)这一领域的联合创始人。该领域致力于理解AI模型内部究竟发生了什么。因此,我们安排他及我们早期的一支团队专注于可解释性研究,因为我们认为这有助于提升模型的安全性与透明度。

英文原文

I’ll give a few examples of this. Early in the history of Anthropic, one of our co-founders, Chris Olah, who I believe you’re interviewing soon, he’s the co-founder of the field of mechanistic interpretability, which is an attempt to understand what’s going on inside AI models. So we had him and one of our early teams focus on this area of interpretability, which we think is good for making models safe and transparent.

Dario Amodei21:48

这一方向在最初三四年里完全没有任何商业应用,至今仍无实际商业应用。目前我们正开展一些早期测试,未来或许终将实现商业化,但这无疑是一条极其漫长的研究路径。而且,我们始终以开源方式推进,并公开分享全部研究成果。我们之所以这么做,是因为相信这是提升模型安全性的有效途径。有趣的是,随着我们持续推进这项工作,其他公司也开始跟进:有些是受我们启发而启动,有些则担心若其他公司纷纷开展此类工作、显得更具责任感,自己若不跟进便会显得缺乏责任感——毕竟没人愿意被视为不负责任的一方。于是他们也采纳了这一方向。当人才加入Anthropic时,可解释性常是吸引他们的关键因素之一;我常对他们说:‘你没去的那些地方,不妨告诉他们你为何选择来这儿。’很快你就会发现,其他公司也陆续成立了各自的可解释性团队。

英文原文

For three or four years that had no commercial application whatsoever. It still doesn’t. Today we’re doing some early betas with it, and probably it will eventually, but this is a very, very long research bed, and one in which we’ve built in public and shared our results publicly. And we did this because we think it’s a way to make models safer. An interesting thing is that as we’ve done this, other companies have started doing it as well. In some cases because they’ve been inspired by it, in some cases because they’re worried that if other companies are doing this, look more responsible, they want to look more responsible too. No one wants to look like the irresponsible actor. And so they adopt this as well. When folks come to Anthropic, interpretability is often a draw, and I tell them, “The other places you didn’t go, tell them why you came here.” And then you see soon that there’s interpretability teams elsewhere as well.

Dario Amodei22:47

某种程度上,这反而削弱了我们的竞争优势,因为‘哦,现在别人也在做了’。但它对整个系统有益,因此我们必须不断开创一些他人尚未充分开展的新方向。其根本目标,是整体抬高‘做正确之事’的重要性。这并非特指我们自身,也非强调某一家‘好人’公司。其他公司同样可以这样做。若它们也加入这场‘竞相向善’的竞赛,那将是最好的消息。其本质在于塑造向上而非向下倾斜的激励机制。

英文原文

And in a way that takes away our competitive advantage, because it’s like, “Oh, now others are doing it as well.” But it’s good for the broader system, and so we have to invent some new thing that we’re doing that others aren’t doing as well. And the hope is to basically bid up the importance of doing the right thing. And it’s not about us in particular. It’s not about having one particular good guy. Other companies can do this as well. If they join the race to do this, that’s the best news ever. It’s about shaping the incentives to point upward instead of shaping the incentives to point downward.

Lex Fridman23:25

我们应明确指出:机制可解释性这一领域,恰恰代表了一种严谨、非空泛、非拍脑袋式的AI安全研究路径——

英文原文

And we should say this example of the field of mechanistic interpretability is just a rigorous non-hand wavy wave doing AI safety-

Dario Amodei23:34

是的。

英文原文

Yes.

Lex Fridman23:34

——或者说,它正朝着这个方向发展。

英文原文

… or it’s tending that way.

Dario Amodei23:36

我们正努力朝此迈进。我的意思是,就我们当前的观察能力而言,仍处于早期阶段;但我惊讶于我们已能如此深入地探查这些系统内部,并理解所见之物。与缩放定律不同——后者仿佛存在某种支配性规律,驱动模型性能不断提升——而在模型内部,模型本身却……并无理由被设计成便于人类理解,对吧?它们的设计初衷是运行、是工作,就像人脑或人体生化系统一样:它们并非为供人类打开盖子、探入内部并加以理解而设计。但我们确实发现——关于这一点,你可以向克里斯作更详尽的了解——当我们打开模型、深入探查时,所发现的内容竟出人意料地引人入胜。

英文原文

Trying to. I mean, I think we’re still early in terms of our ability to see things, but I’ve been surprised at how much we’ve been able to look inside these systems and understand what we see. Unlike with the scaling laws where it feels like there’s some law that’s driving these models to perform better, on the inside, the models aren’t… There’s no reason why they should be designed for us to understand them, right? They’re designed to operate, they’re designed to work. Just like the human brain or human biochemistry. They’re not designed for a human to open up the hatch, look inside and understand them. But we have found, and you can talk in much more detail about this to Chris, that when we open them up, when we do look inside them, we find things that are surprisingly interesting.

Lex Fridman24:20

作为附带效果,你还能从中领略这些模型之美;你能借助MEC与TERP这类方法论,探索大型神经网络所蕴含的精妙之美。

英文原文

And as a side effect, you also get to see the beauty of these models. You get to explore the beautiful nature of large neural networks through the MEC and TERP kind of methodology.

Dario Amodei24:29

我对其中的简洁性深感震撼,例如‘归纳头’(induction heads);我也对能利用稀疏自编码器在神经网络中定位特定方向感到惊叹,而这些方向竟对应着极为清晰的概念。

英文原文

I’m amazed at how clean it’s been. I’m amazed at things like induction heads. I’m amazed at things like that we can use sparse auto-encoders to find these directions within the networks, and that the directions correspond to these very clear concepts.

Dario Amodei24:49

我们曾在‘金门大桥版Claude’实验中部分展示了这一点。该实验中,我们在某一层神经网络内发现了一个与‘金门大桥’相对应的方向,随后仅将该方向的激活程度调高。我们以此模型发布了一个演示版本——略带玩笑性质,仅上线数日,但生动展现了我们所开发的方法。你可以向该模型提问任何问题:例如‘你今天过得如何?’——由于该特征被持续激活,无论你问什么,回答都会关联到金门大桥。例如它会说:‘我感觉放松而开阔,正如金门大桥的拱形结构一般……’

英文原文

We demonstrated this a bit with the Golden Gate Bridge Claude. So this was an experiment where we found a direction inside one of the neural networks layers that corresponded to the Golden Gate Bridge. And we just turned that way up. And so we released this model as a demo, it was kind of half a joke, for a couple days, but it was illustrative of the method we developed. And you could take the model, you could ask it about anything. It would be like you could say, “How was your day?” And anything you asked, because this feature was activated, it would connect to the Golden Gate Bridge. So it would say, I’m feeling relaxed and expansive, much like the arches of the Golden Gate Bridge, or-

Lex Fridman25:31

它会巧妙地将话题转向金门大桥并自然融入其中。同时,它对金门大桥的专注还透着一丝忧郁。我想人们很快便爱上了它,甚至在它下线(大概仅一天后)后便开始怀念。

英文原文

It would masterfully change topic to the Golden Gate Bridge and integrate it. There was also a sadness to the focus it had on the Golden Gate Bridge. I think people quickly fell in love with it, I think. So people already miss it, because it was taken down, I think after a day.

Dario Amodei25:45

不知为何,这类对模型的行为干预——即调整其行为模式——竟在情感层面使其显得比其他任何版本的模型都更像人类。

英文原文

Somehow these interventions on the model, where you kind of adjust its behavior, somehow emotionally made it seem more human than any other version of the model.

Lex Fridman25:56

它拥有鲜明的个性与强烈的身份认同。

英文原文

It’s a strong personality, strong identity.

Dario Amodei25:58

它个性鲜明,且带有某种近乎偏执的兴趣点。我们每个人都能想到某个对此类事物极度痴迷的人。因此,它确实在某种程度上让人感觉更富人性。

英文原文

It has a strong personality. It has these kind of obsessive interests. We can all think of someone who’s obsessed with something. So it does make it feel somehow a bit more human.

Lex Fridman26:08

让我们谈谈当下,谈谈Claude。今年发生了许多事情:3月,Claude 3 Opus、Sonnet与Haiku相继发布;7月推出Claude 3.5 Sonnet,最新更新版刚刚发布;此外,Claude 3.5 Haiku也已发布。那么,能否请您解释一下Opus、Sonnet与Haiku之间的区别,以及我们应如何理解这些不同版本?

英文原文

Let’s talk about the present. Let’s talk about Claude. So this year, a lot has happened. In March. Claude 3 Opus, Sonnet, Haiku were released. Then Claude 3.5 Sonnet in July, with an updated version just now released. And then also Claude 3.5 Haiku was released. Okay. Can you explain the difference between Opus, Sonnet and Haiku, and how we should think about the different versions?

Dario Amodei26:34

是的,那我们回到三月份,也就是我们首次发布这三款模型的时候。当时我们的想法是:不同公司会推出大小不一、优劣各异的大模型。我们认为市场既需要一款真正强大的模型——它可能稍慢一些,价格也更高;也需要快速、廉价的模型——在保证速度和成本的前提下,尽可能聪明。当你需要执行某种难度较高的分析任务时,比如我要编写代码、头脑风暴创意点子,或者进行创意写作,我就会选择这款真正强大的模型。

英文原文

Yeah, so let’s go back to March when we first released these three models. So our thinking was different companies produce large and small models, better and worse models. We felt that there was demand, both for a really powerful model, and that might be a little bit slower that you’d have to pay more for, and also for fast cheap models that are as smart as they can be for how fast and cheap. Whenever you want to do some kind of difficult analysis, like if I want to write code for instance, or I want to brainstorm ideas or I want to do creative writing, I want the really powerful model.

Dario Amodei27:15

但另一方面,在商业应用中还有大量实际场景,例如:我在浏览一个网站、申报个人所得税、或向法律顾问咨询并分析一份合同。此外,还有很多公司只是希望在自己的集成开发环境(IDE)中实现自动补全之类的功能。对于所有这些用途,你都希望模型响应迅速,并能被大规模广泛使用。因此,我们希望全面覆盖这一整条需求光谱。最终,我们采用了‘诗歌’这一主题来命名:那么,最短小的诗是什么?是俳句(Haiku)。俳句代表的是体积最小、响应最快、成本最低的模型;而当时这款模型的智能程度,对于其速度与成本而言,确实令人惊喜。

英文原文

But then there’s a lot of practical applications in a business sense where it’s like I’m interacting with a website, I am doing my taxes, or I’m talking to a legal advisor and I want to analyze a contract. Or we have plenty of companies that are just like, I want to do auto-complete on my IDE or something. And for all of those things, you want to act fast and you want to use the model very broadly. So we wanted to serve that whole spectrum of needs. So we ended up with this kind of poetry theme. And so what’s a really short poem? It’s a haiku. Haiku is the small, fast, cheap model that was at the time, was really surprisingly intelligent for how fast and cheap it was.

Dario Amodei28:03

十四行诗(Sonnet)是一种中等篇幅的诗体,通常写上几段文字。因此,Sonnet 就是我们定位为中等规模的模型:它比 Haiku 更聪明,但也略慢一些,价格也略高一些。而 Opus(杰作)则如‘鸿篇巨制’(Magnum Opus)一般,代表宏大的作品;Opus 当时就是规模最大、最聪明的模型。这便是最初命名背后的思路。

英文原文

Sonnet is a medium-sized poem, write a couple paragraphs. And so Sonnet was the middle model. It is smarter but also a little bit slower, a little bit more expensive. And Opus, like a Magnum Opus is a large work, Opus was the largest, smartest model at the time. So that was the original kind of thinking behind it.

Dario Amodei28:24

随后我们的思路是:‘每一代新模型都应推动这条权衡曲线发生位移。’ 举例来说,当我们发布 Sonnet 3.5 时,它的成本和速度大致与 Sonnet 3 相当,但其智能水平已提升至超越初代 Opus 3 的程度——尤其在编程方面,整体表现亦更优。目前我们已公布了 Haiku 3.5 的测试结果。据我所知,这款最新推出的最小模型 Haiku 3.5,其能力已基本达到旧版最大模型 Opus 3 的水平。因此,我们的核心目标正是持续推动这条曲线前移,而未来某一天,自然也会出现 Opus 3.5。

英文原文

And our thinking then was, “Well, each new generation of models should shift that trade- off curve.” So when we released Sonnet 3.5, it has roughly the same cost and speed as the Sonnet 3 model, but it increased its intelligence to the point where it was smarter than the original Opus 3 model. Especially for code, but also just in general. And so now we’ve shown results for Haiku 3.5. And I believe Haiku 3.5, the smallest new model, is about as good as Opus 3, the largest old model. So basically the aim here is to shift the curve and then at some point there’s going to be an Opus 3.5.

Dario Amodei29:13

如今,每一代新模型都有其独特之处:它们使用新的训练数据,其‘个性’也会以我们试图引导、却无法完全掌控的方式发生变化。因此,永远不存在一种‘仅改变智能水平’的精确等价关系。我们始终致力于同步改进其他各项指标,而有些变化甚至在我们未察觉或未测量的情况下就已发生。所以,这本质上是一门极不精确的科学。在许多方面,这些模型的风格与个性,与其说是一门科学,不如说更像一门艺术。

英文原文

Now every new generation of models has its own thing. They use new data, their personality changes in ways that we try to steer but are not fully able to steer. And so there’s never quite that exact equivalence, where the only thing you’re changing is intelligence. We always try and improve other things and some things change without us knowing or measuring. So it’s very much an inexact science. In many ways, the manner and personality of these models is more an art than it is a science.

Lex Fridman29:44

那么,从 Claude Opus 3.0 到 3.5 之间相隔的时间跨度,其原因何在?如果方便的话,能否具体说明一下?

英文原文

So what is the reason for the span of time between say, Claude Opus 3.0 and 3.5? What takes that time, if you can speak to it?

Dario Amodei29:58

是的,整个流程包含多个环节。首先是预训练(pre-training),即常规的语言模型训练阶段,耗时非常长。当前阶段,我们动辄使用数万乃至数万以上的 GPU 或 TPU 进行训练;我们也会采用不同平台,但核心都是各类加速芯片,训练周期往往长达数月。

英文原文

Yeah, so there’s different processes. There’s pre-training, which is just kind of the normal language model training. And that takes a very long time. That uses, these days, tens of thousands, sometimes many tens of thousands of GPUs or TPUs or training them, or we use different platforms, but accelerator chips, often training for months.

Dario Amodei30:26

接着是后训练(post-training)阶段,我们会开展基于人类反馈的强化学习(RLHF),以及其他类型的强化学习。这一阶段如今正变得越来越庞大,且远非一门精确科学;要调优到位往往需投入大量人力。随后,我们会邀请部分早期合作伙伴对模型进行试用评估,检验其实际效果;同时,我们还会在内部及外部对其安全性展开严格测试,尤其关注灾难性风险与自主性风险。我们依据自身的《负责任扩展政策》(Responsible Scaling Policy)开展内部安全测试,这方面我也可以进一步详细说明。

英文原文

There’s then a kind of post-training phase where we do reinforcement learning from human feedback as well as other kinds of reinforcement learning. That phase is getting larger and larger now, and often that’s less of an exact science. It often takes effort to get it right. Models are then tested with some of our early partners to see how good they are, and they’re then tested, both internally and externally, for their safety, particularly for catastrophic and autonomy risks. So we do internal testing according to our responsible scaling policy, which I could talk more about that in detail.

Dario Amodei31:06

此外,我们还与美国和英国的人工智能安全研究所(AI Safety Institute),以及特定领域内的其他第三方评测机构达成协议,对模型开展所谓‘CBRN 风险’(即化学、生物、放射性与核风险)的专项评估。我们并不认为当前模型已实质性具备此类风险,但每推出一款新模型,我们都希望评估其是否正逐步接近某些更具危险性的能力边界。以上便是主要流程阶段;此外,还需额外时间完成模型推理(inference)的工程化适配,并将其正式上线至 API 接口。因此,让一款模型真正落地可用,本身便涉及大量环节。当然,我们始终致力于尽可能优化和精简整个流程。

英文原文

And then we have an agreement with the US and the UK AI Safety Institute, as well as other third-party testers in specific domains, to test the models for what are called CBRN risks, chemical, biological, radiological, and nuclear. We don’t think that models pose these risks seriously yet, but every new model we want to evaluate to see if we’re starting to get close to some of these more dangerous capabilities. So those are the phases, and then it just takes some time to get the model working in terms of inference and launching it in the API. So there’s just a lot of steps to actually making a model work. And of course, we’re always trying to make the processes as streamlined as possible.

Dario Amodei31:55

我们希望安全测试既严谨,又高效自动化——在不牺牲严谨性的前提下,尽可能加快测试节奏。预训练与后训练流程同样如此。这就像制造任何复杂产品一样,譬如造飞机:既要确保绝对安全,又要让整个生产流程高度顺畅。我认为,这种安全与效率之间的创造性张力,恰恰是推动模型成功落地的关键所在。

英文原文

We want our safety testing to be rigorous, but we want it to be rigorous and to be automatic, to happen as fast as it can, without compromising on rigor. Same with our pre-training process and our post-training process. So it’s just building anything else. It’s just like building airplanes. You want to make them safe, but you want to make the process streamlined. And I think the creative tension between those is an important thing in making the models work.

Lex Fridman32:20

坊间传闻——我记不清是谁说的了——Anthropic 公司的工程工具链(tooling)非常出色。因此,当前挑战中很大一部分,很可能落在软件工程层面:即构建一套高效、低摩擦的工具体系,以实现与底层基础设施的顺畅交互。

英文原文

Yeah, rumor on the street, I forget who was saying that, Anthropic has really good tooling. So probably a lot of the challenge here is, on the software engineering side, is to build the tooling to have a efficient, low-friction interaction with the infrastructure.

Dario Amodei32:36

你会惊讶地发现,构建这类模型所面临的诸多挑战,最终大多归结于软件工程与性能工程。从外部看,人们或许会想:‘天啊,我们有了个尤里卡式突破!’——就像电影里演的那样,充满科学感:‘我们发现了它,我们搞懂了它。’但事实上,一切成果,哪怕再惊人的发现,几乎总是取决于无数细节,而且往往是极其、极其枯燥的细节。至于我们是否比其他公司拥有更出色的工具链,我无法断言。毕竟我没在那些公司工作过——至少近期没有。但可以肯定的是,这是我们长期高度重视的领域。

英文原文

You would be surprised how much of the challenges of building these models comes down to software engineering, performance engineering. From the outside, you might think, “Oh man, we had this Eureka breakthrough.” You know, this movie with the science. “We discovered it, we figured it out.” But I think all things, even incredible discoveries, they almost always come down to the details. And often super, super boring details. I can’t speak to whether we have better tooling than other companies. I mean, haven’t been at those other companies, at least not recently, but it’s certainly something we give a lot of attention to.

Lex Fridman33:18

我不确定您是否方便透露:从 Claude 3 到 Claude 3.5,是否存在额外的预训练环节?还是主要聚焦于后训练?毕竟性能提升幅度相当显著。

英文原文

I don’t know if you can say, but from Claude 3 to Claude 3.5, is there any extra pre-training going on, or is it mostly focused on the post-training? There’s been leaps in performance.

Dario Amodei33:29

是的,我认为在任一阶段,我们天然都会聚焦于全方位同步提升。比如,不同团队各自负责不同环节,每个团队都在努力优化自己所承担的‘接力赛’中的那一棒。因此,当推出一款新模型时,我们自然会将所有这些进步成果一次性整合进去。

英文原文

Yeah, I think at any given stage, we’re focused on improving everything at once. Just naturally. Like, there are different teams. Each team makes progress in a particular area, in making their particular segment of the relay race better. And it’s just natural that when we make a new model, we put all of these things in at once.

Lex Fridman33:50

那么,您所积累的 RLHF 偏好数据,是否能在新模型训练过程中复用?

英文原文

So the data you have, the preference data you get from RLHF, is there ways to apply it to newer models as it get trained up?

Dario Amodei34:00

是的。旧模型产生的偏好数据有时也会用于训练新模型,尽管显然在新模型上直接采集并训练的效果会更好。请注意,我们采用‘宪法式人工智能’(Constitutional AI)方法,因此不仅依赖偏好数据;还存在另一类后训练流程,即让模型针对自身进行训练。每天我们都会引入新型的、基于模型自反馈的后训练方式。因此,后训练绝不仅限于 RLHF,还包括大量其他方法。我认为,后训练正变得日益复杂与精密。

英文原文

Yeah. Preference data from old models sometimes gets used for new models, although of course it performs somewhat better when it’s trained on the new models. Note that we have this constitutional AI method such that we don’t only use preference data, there’s also a post-training process where we train the model against itself. And there’s new types of post-training the model against itself that are used every day. So it’s not just RLHF, a bunch of other methods as well. Post-training, I think, is becoming more and more sophisticated.

Lex Fridman34:30

那么,Sonnet 3.5 在编程能力方面为何实现了如此显著的跃升?至少在编程侧表现突出。也许这正是讨论基准测试(benchmarks)的好时机:所谓‘变强’究竟意味着什么?仅仅是分数提高了?我自己也写程序,而且热爱编程;我日常借助 Cursor 工具使用 Claude 3.5 辅助编程。从实际体验和坊间反馈来看,它在编程方面确实变得更聪明了。那么,到底需要做些什么,才能让它变得更聪明?

英文原文

Well, what explains the big leap in performance for the new Sonnet 3.5, I mean, at least in the programming side? And maybe this is a good place to talk about benchmarks. What does it mean to get better? Just the number went up, but I program, but I also love programming, and I Claude 3.5 through Cursor is what I use to assist me in programming. And there was, at least experientially, anecdotally, it’s gotten smarter at programming. So what does it take to get it smarter?

Dario Amodei34:30

我们——

英文原文

We-

Lex Fridman35:00

那么,到底需要做些什么,才能让它变得更聪明?

英文原文

So what does it take to get it smarter?

Dario Amodei35:03

我们自己也观察到了这一点。顺便提一句,Anthropic 内部有几位非常资深的工程师,此前无论是我们自己还是其他公司发布的各类代码模型,对他们而言都基本无实用价值。他们曾表示:‘也许这对初学者有点用,但对我毫无帮助。’然而,Sonnet 3.5(最初的版本)首次让他们感叹:‘天啊,它帮我解决了一个原本需要耗费数小时才能完成的问题!这是第一款真正为我节省了时间的模型。’

英文原文

We observe that as well. By the way, there were a couple very strong engineers here at Anthropic, who all previous code models, both produced by us and produced by all the other companies, hadn’t really been useful to them. They said, “Maybe this is useful to a beginner. It’s not useful to me.” But Sonnet 3.5, the original one for the first time, they said, “Oh, my God, this helped me with something that it would’ve taken me hours to do. This is the first model that’s actually saved me time.”

Dario Amodei35:31

所以,再次强调,水位线正在上升。接着,我认为新版 Sonnet 表现得甚至更出色了。就其所需能力而言,我直接说吧:整体上都有提升——体现在预训练阶段、后训练阶段,以及我们开展的各项评估中。我们自己也观察到了这一点。如果我们深入基准测试的细节,SWE-bench 实际上是……既然你是一名程序员,那你肯定熟悉‘拉取请求(pull requests)’,而拉取请求本身就像一种工作原子单元。你可以理解为:我正在实现某一项具体功能。

英文原文

So again, the water line is rising. And then I think the new Sonnet has been even better. In terms of what it takes, I’ll just say it’s been across the board. It’s in the pre-training, it’s in the post-training, it’s in various evaluations that we do. We’ve observed this as well. And if we go into the details of the benchmark, so SWE-bench is basically… Since you’re a programmer, you’ll be familiar with pull requests, and just pull requests, they’re like a sort of atomic unit of work. You could say I’m implementing one thing.

Dario Amodei36:12

因此,SWE-bench 实际上为你提供了一个真实场景:代码库处于当前状态,而我需要根据自然语言描述来实现某项功能。我们内部也有类似的基准测试,用以衡量相同的能力;测试方式是:‘完全放开模型限制,允许它自由执行任意操作、运行任意代码、编辑任意内容——它完成这些任务的能力究竟如何?’ 正是这一基准测试的结果,从‘仅能在 3% 的情况下完成任务’提升到了‘约 50% 的情况下能完成任务’。

英文原文

So SWE-bench actually gives you a real world situation where the code base is in a current state and I’m trying to implement something that’s described in language. We have internal benchmarks where we measure the same thing and you say, “Just give the model free rein to do anything, run anything, edit anything. How well is it able to complete these tasks?” And it’s that benchmark that’s gone from “it can do it 3% of the time” to “it can do it about 50% of the time.”

Dario Amodei36:43

因此,我确实相信基准测试分数是可以提升的;但我觉得,如果我们在不针对该特定基准过度训练或‘刷分’的前提下,真正达到 100% 的通过率,那很可能代表编程能力出现了真实且重大的提升。我推测,若能达到 90% 或 95% 的通过率,则意味着模型已具备自主完成相当大比例软件工程任务的能力。

英文原文

So I actually do believe that you can gain benchmarks, but I think if we get to 100% on that benchmark in a way that isn’t over-trained or game for that particular benchmark, probably represents a real and serious increase in programming ability. And I would suspect that if we can get to 90, 95% that it will represent ability to autonomously do a significant fraction of software engineering tasks.

Lex Fridman37:13

好吧,问个荒谬的时间线问题:Claude Opus 3.5 什么时候发布?

英文原文

Well, ridiculous timeline question. When is Claude Opus 3.5 coming up?

Dario Amodei37:19

我不会给出确切日期,但据我们所知,计划仍是要推出 Claude 3.5 Opus。

英文原文

Not giving you an exact date, but as far as we know, the plan is still to have a Claude 3.5 Opus.

Lex Fridman37:28

我们会在《GTA 6》之前拿到它吗?还是不会?

英文原文

Are we going to get it before GTA 6 or no?

Dario Amodei37:30

就像《永远的毁灭公爵》(Duke Nukem Forever)那样?

英文原文

Like Duke Nukem Forever?

Lex Fridman37:30

《永远的毁灭公爵》,没错。

英文原文

Duke Nukem. Right.

Dario Amodei37:32

那款游戏叫什么来着?有款游戏被推迟了整整 15 年。

英文原文

What was that game? There was some game that was delayed 15 years.

Lex Fridman37:32

没错。

英文原文

That’s right.

Dario Amodei37:34

是《永远的毁灭公爵》吗?

英文原文

Was that Duke Nukem Forever?

Lex Fridman37:36

是的。而且我觉得《GTA》现在刚放出预告片而已。

英文原文

Yeah. And I think GTA is now just releasing trailers.

Dario Amodei37:39

距离我们首次发布 Sonnet 才过去三个月。

英文原文

It’s only been three months since we released the first Sonnet.

Lex Fridman37:42

是啊,发布节奏实在惊人。

英文原文

Yeah, it’s the incredible pace of release.

Dario Amodei37:45

这恰恰说明了研发节奏之快,以及外界对产品发布时间的预期之高。

英文原文

It just tells you about the pace, the expectations for when things are going to come out.

Lex Fridman37:49

那么关于 4.0 呢?随着模型越来越大,你们如何看待版本号命名问题?另外,就版本命名本身而言,为什么 Sonnet 3.5 更新时要加上发布日期?为什么不叫 Sonnet 3.6——很多人其实已经在这么叫了?

英文原文

So what about 4.0? So how do you think, as these models get bigger and bigger, about versioning and also just versioning in general, why Sonnet 3.5 updated with the date? Why not Sonnet 3.6, which a lot of people are calling it?

Dario Amodei38:06

命名实际上在这里是个挺有意思的问题,对吧?因为我想一年前,大多数模型还主要依赖预训练。所以你可以从头开始规划:‘好,我们将推出不同尺寸的模型,全部一起训练,形成一套命名体系;然后往里面加入一些新魔法,再推出下一代。’

英文原文

Naming is actually an interesting challenge here, right? Because I think a year ago, most of the model was pre-training. And so you could start from the beginning and just say, “Okay, we’re going to have models of different sizes. We’re going to train them all together and we’ll have a family of naming schemes and then we’ll put some new magic into them and then we’ll have the next generation.”

Dario Amodei38:26

麻烦一开始就会出现——有些模型的训练耗时远超其他模型,这已经略微打乱了时间安排。而当你在预训练环节取得重大改进时,你突然意识到:‘哦,我能做出更好的预训练模型了。’ 这类改进往往不需要太长时间,但显然,新模型与旧模型在尺寸和结构上完全一致。所以,我认为上述两点,再加上发布时间上的各种不确定性,使得任何你设计出来的命名方案,现实情况都很容易让这套方案失效,对吧?它总会突破这套方案的框架。

英文原文

The trouble starts already when some of them take a lot longer than others to train. That already messes up your time a little bit. But as you make big improvement in pre-training, then you suddenly notice, “Oh, I can make better pre-train model.” And that doesn’t take very long to do, but clearly it has the same size and shape of previous models. So I think those two together as well as the timing issues. Any kind of scheme you come up with, the reality tends to frustrate that scheme, right? It tends to break out of the scheme.

Dario Amodei39:04

这不像软件开发,你不能简单地说‘这是 3.7 版,那是 3.8 版’。不,我们的模型各有不同的权衡取舍:你可以调整模型中的某些部分,也可以调整其他部分;有些模型推理更快、有些更慢;有些必须更昂贵,有些则必须更便宜。所以我认为所有公司都在为此挣扎。就命名而言,我们当初推出 Haiku、Sonnet 和 Opus 时,处境其实相当有利。

英文原文

It’s not like software where you can say, “Oh, this is 3.7, this is 3.8.” No, you have models with different trade-offs. You can change some things in your models, you can change other things. Some are faster and slower at inference. Some have to be more expensive, some have to be less expensive. And so I think all the companies have struggled with this. I think we were in a good position in terms of naming when we had Haiku, Sonnet and Opus.

Lex Fridman39:31

那真是个极好的、极棒的开端。

英文原文

It was great, great start.

Dario Amodei39:32

我们正努力维持这套命名体系,但它并不完美,因此我们会尝试回归简洁性。但就这个领域的本质而言,我觉得目前还没有人真正搞懂命名这件事——它似乎与常规软件开发遵循的是截然不同的范式,因此没有任何一家公司在命名上做到尽善尽美。相比起训练模型这项宏大的科学工程,命名看似微不足道,但我们却为此付出了出乎意料的巨大精力。

英文原文

We’re trying to maintain it, but it’s not perfect, so we’ll try and get back to the simplicity. But just the nature of the field, I feel like no one’s figured out naming. It’s somehow a different paradigm from normal software and so none of the companies have been perfect at it. It’s something we struggle with surprisingly much relative to how trivial it is for the grand science of training the models.

Lex Fridman40:03

因此,从用户角度看,更新后的 Sonnet 3.5 与此前 2024 年 6 月发布的 Sonnet 3.5 在用户体验上明显不同。若能设计出某种标签体系来体现这种差异,那就再好不过了。因为人们口中的‘Sonnet 3.5’如今已指代两个不同版本,那么当存在明确性能提升时,该如何区分旧版与新版?这使得围绕它的讨论本身就变得十分困难。

英文原文

So from the user side, the user experience of the updated Sonnet 3.5 is just different than the previous June 2024 Sonnet 3.5. It would be nice to come up with some kind of labeling that embodies that. Because people talk about Sonnet 3.5, but now there’s a different one. And so how do you refer to the previous one and the new one when there’s a distinct improvement? It just makes conversation about it just challenging.

Dario Amodei40:34

是的,是的。我确实认为,模型存在大量未被基准测试所反映的属性,这一点毫无疑问,大家也都认同。而且,并非所有属性都属于‘能力’范畴:模型可以表现得礼貌或生硬,可以反应迅速或主动提问,可以给人温暖亲切的感觉,也可以显得冷漠疏离;它可以枯燥乏味,也可以极具个性——比如‘金门大桥版 Claude’(Golden Gate Claude)就是如此。

英文原文

Yeah, yeah. I definitely think this question of there are lots of properties of the models that are not reflected in the benchmarks. I think that’s definitely the case and everyone agrees. And not all of them are capabilities. Models can be polite or brusque, they can be very reactive or they can ask you questions. They can have what feels like a warm personality or a cold personality. They can be boring or they can be very distinctive like Golden Gate Claude was.

Dario Amodei41:10

我们有一个专门团队专注于这方面工作,我想我们称之为‘Claude 角色(Claude character)’。阿曼达(Amanda)领导该团队,稍后我们也会就此与你展开交流。但目前这仍是一门极不精确的科学;我们常常发现,模型会表现出一些我们事先并未察觉的特性。事实是,你可能与一个模型对话一万次,却仍无法观察到它的某些行为——这跟人类其实很像,对吧?

英文原文

And we have a whole team focused on, I think we call it Claude character. Amanda leads that team and we’ll talk to you about that, but it’s still a very inexact science and often we find that models have properties that we’re not aware of. The fact of the matter is that you can talk to a model 10,000 times and there are some behaviors you might not see just like with a human, right?

Dario Amodei41:36

我可能认识某个人几个月,却仍不知道他拥有某项技能,或不了解他性格中某个特定侧面。因此,我认为我们必须逐渐适应这种认知。我们始终在寻找更优的方法来测试模型,以验证这些能力,同时也决定哪些人格特质是我们希望模型具备的、哪些又是我们不希望它拥有的。而这个规范性问题本身,也极其有趣。

英文原文

I can know someone for a few months and not know that they have a certain skill or not know that there’s a certain side to them. And so I think we just have to get used to this idea. And we’re always looking for better ways of testing our models to demonstrate these capabilities and also to decide which are the personality properties we want models to have and which we don’t want to have. That itself, the normative question, is also super interesting.

Lex Fridman42:02

我得向你提一个来自 Reddit 的问题。

英文原文

I got to ask you a question from Reddit.

Dario Amodei42:04

来自 Reddit?哎呀,天哪。

英文原文

From Reddit? Oh, boy.

Lex Fridman42:07

这现象对我而言至少非常有趣——它是一种引人入胜的心理社会现象:许多人报告称,Claude 在他们使用过程中‘变笨了’。因此问题是:用户关于 Claude 3.5 Sonnet ‘变笨’的抱怨是否站得住脚?这些轶事性报告究竟只是一种社会心理现象,还是真存在 Claude 变得更‘笨’的案例?

英文原文

There’s just this fascinating, to me at least, it’s a psychological social phenomenon where people report that Claude has gotten dumber for them over time. And so the question is, does the user complaint about the dumbing down of Claude 3.5 Sonnet hold any water? So are these anecdotal reports a kind of social phenomena or is there any cases where Claude would get dumber?

Dario Amodei42:33

事实上,这种情况并不适用。这并不仅限于 Claude。我相信,每家大型公司推出的每个基础模型,都曾遭遇过类似投诉。人们曾这样评价 GPT-4,也曾这样评价 GPT-4 Turbo。原因有二:第一,模型的实际权重(即模型真正的‘大脑’)除非我们发布全新模型,否则绝不会改变;从实践角度出发,随机替换模型的新版本根本毫无意义。

英文原文

So this actually doesn’t apply. This isn’t just about Claude. I believe I’ve seen these complaints for every foundation model produced by a major company. People said this about GPT-4, they said it about GPT-4 Turbo. So a couple things. One, the actual weights of the model, the actual brain of the model, that does not change unless we introduce a new model. There are just a number of reasons why it would not make sense practically to be randomly substituting in new versions of the model.

Dario Amodei43:09

从推理层面看这很难实现,而且实际更改模型权重所引发的连锁后果也极难全面掌控。举例来说,假设你想对模型进行微调,比如让它少说‘当然(certainly)’——而旧版 Sonnet 确实常这么说——结果你往往会同时改变一百个其他方面。因此,我们有一整套流程来处理模型修改,包括大量测试及面向早期客户的用户测试。

英文原文

It’s difficult from an inference perspective and it’s actually hard to control all the consequences of changing the weights of the model. Let’s say you wanted to fine-tune the model, I don’t know, to say “certainly” less, which an old version of Sonnet used to do. You actually end up changing 100 things as well. So we have a whole process for it and we have a whole process for modifying the model. We do a bunch of testing on it. We do a bunch of user testing in early customers.

Dario Amodei43:36

因此,我们从未在未告知任何人的情况下擅自更改过模型权重;而且,在当前架构下,这么做也完全不合逻辑。不过,我们偶尔确实会做两件事:一是有时会开展 A/B 测试,但这类测试通常紧邻新模型发布前后,且持续时间极短。

英文原文

So we both have never changed the weights of the model without telling anyone. And certainly, in the current setup, it would not make sense to do that. Now, there are a couple things that we do occasionally do. One is sometimes we run A/B tests, but those are typically very close to when a model is being released and for a very small fraction of time.

Dario Amodei44:01

所以,在新版 Sonnet 3.5 发布前一天,我同意:我们本该给它起个更好的名字——目前这种叫法实在拗口。当时有人评论说模型‘进步巨大’,原因正是他们在那一两天内偶然参与了 A/B 测试。另一件事是:系统提示词(system prompt)偶尔会发生变化。系统提示词确实会产生一定影响,但几乎不可能导致模型‘变笨’,更不可能让它变得更迟钝。

英文原文

So the day before the new Sonnet 3.5, I agree we should have had a better name. It’s clunky to refer to it. There were some comments from people that it’s gotten a lot better and that’s because a fraction we’re exposed to an A/B test for those one or two days. The other is that occasionally the system prompt will change. The system prompt can have some effects, although it’s unlikely to dumb down models, it’s unlikely to make them dumber.

Dario Amodei44:32

我们已经看到,虽然我为力求全面而列出的这两件事发生得相当罕见,但对我们及其他模型公司而言,关于模型变更的投诉却持续不断——比如‘模型不擅长这个’‘模型变得更受审查了’‘模型被弱化了’。这些投诉始终存在,因此我不想说人们是在凭空想象之类的话,但事实上,模型在绝大多数情况下并未发生变化。如果让我提出一个理论,我认为它其实与我之前提到的一点有关,即模型极为复杂,具有诸多不同维度。因此,当我向模型提问时,若我说‘执行任务X’, versus ‘你能执行任务X吗?’,模型的回应可能截然不同。所以,在与模型交互的方式上,存在大量细微差异,而这些差异可能导致结果大相径庭。

英文原文

And we’ve seen that while these two things, which I’m listing to be very complete, happened quite infrequently, the complaints for us and for other model companies about the model change, the model isn’t good at this, the model got more censored, the model was dumbed down. Those complaints are constant and so I don’t want to say people are imagining it or anything, but the models are, for the most part, not changing. If I were to offer a theory, I think it actually relates to one of the things I said before, which is that models are very complex and have many aspects to them. And so often, if I ask the model a question, if I’m like, “Do task X” versus, “Can you do task X?” the model might respond in different ways. And so there are all kinds of subtle things that you can change about the way you interact with the model that can give you very different results.

Dario Amodei45:33

需要明确的是,这本身恰恰暴露了我们及其他模型提供商的一项缺陷:模型往往对措辞上的微小变动异常敏感。这也再次印证,关于这些模型如何运作的科学目前仍极不成熟。因此,假如我某天晚上入睡之前以某种方式与模型对话,而次日稍一调整我的提问措辞,就可能得到完全不同的结果。

英文原文

To be clear, this itself is like a failing by us and by the other model providers that the models are just often sensitive to small changes in wording. It’s yet another way in which the science of how these models work is very poorly developed. And so if I go to sleep one night and I was talking to the model in a certain way and I slightly changed the phrasing of how I talk to the model, I could get different results.

Dario Amodei45:58

所以这是一种可能的解释。另一点是,唉,这类现象本身就极难量化。极难量化。我认为,每当有新模型发布时,人们都会异常兴奋;但随着时间推移,他们又会越来越清楚地意识到其局限性。这或许构成了另一种效应。但总而言之,除少数非常狭窄的情形外,模型实际上并未发生变化——我绕了这么大一圈,想表达的正是这一点。

英文原文

So that’s one possible way. The other thing is, man, it’s just hard to quantify this stuff. It’s hard to quantify this stuff. I think people are very excited by new models when they come out and then as time goes on, they become very aware of their limitations. So that may be another effect, but that’s all a very long-winded way of saying for the most part, with some fairly narrow exceptions, the models are not changing.

Lex Fridman46:22

我认为其中存在一种心理效应:你只是逐渐习惯了它,基准线随之提高。比如,最早在飞机上体验到Wi-Fi的人,会觉得那简直神奇、不可思议。

英文原文

I think there is a psychological effect. You just start getting used to it, the baseline raises. When people who have first gotten Wi-Fi on airplanes, it’s amazing, magic.

Dario Amodei46:32

太神奇了,没错。

英文原文

It’s amazing. Yeah.

Lex Fridman46:32

然后你开始……

英文原文

And then you start-

Dario Amodei46:33

而现在我却会想:‘这玩意儿根本没法用,简直是一堆垃圾。’

英文原文

And now I’m like, “I can’t get this thing to work. This is such a piece of crap.”

Lex Fridman46:36

完全正确。因此,很容易产生这样一种阴谋论:‘他们在让Wi-Fi越来越慢。’这个问题我之后可能会跟阿曼达更深入地探讨。不过,另有一条来自Reddit的提问:‘Claude何时才会停止扮演我那位纯粹的触手系祖母,强行将它的道德世界观强加于我这位付费用户身上?另外,让Claude过度道歉的心理学动因是什么?’这条反馈描述了一种用户体验,从另一个角度折射出用户的挫败感。它涉及角色设定[此处音频不清,00:47:06]。

英文原文

Exactly. So it’s easy to have the conspiracy theory of, “They’re making Wi-Fi slower and slower.” This is probably something I’ll talk to Amanda much more about, but another Reddit question, “When will Claude stop trying to be my pure tentacle grandmother imposing its moral worldview on me as a paying customer? And also, what is the psychology behind making Claude overly apologetic?” So this reports about the experience, a different angle on the frustration. It has to do with the character [inaudible 00:47:06].

Dario Amodei47:06

是的,关于这个问题,首先我想指出两点。第一点是,人们在Reddit、Twitter(或X平台)等社交媒体上所发表的言论,与其实际统计意义上用户真正关心、并驱动其使用模型的因素之间,存在着巨大的分布偏移。

英文原文

Yeah, so a couple points on this first. One is things that people say on Reddit and Twitter or X or whatever it is, there’s actually a huge distribution shift between the stuff that people complain loudly about on social media and what actually statistically users care about and that drives people to use the models.

Dario Amodei47:27

人们对诸如‘模型无法完整写出全部代码’或‘模型在编程方面本可表现得更好,却未能达到预期水平’等问题感到沮丧——尽管它已是全球范围内编程能力最强的模型。我认为,大多数抱怨都集中于这类问题;当然,也确实存在一个声音响亮的少数群体,他们对模型拒绝本不该拒绝的请求、过度道歉,或表现出这些令人厌烦的语言习惯而深感困扰。

英文原文

People are frustrated with things like the model not writing out all the code or the model just not being as good at code as it could be, even though it’s the best model in the world on code. I think the majority of things are about that, but certainly a vocal minority raise these concerns, are frustrated by the model refusing things that it shouldn’t refuse or apologizing too much or just having these annoying verbal tics.

Dario Amodei47:59

第二点需特别说明——我想把这点讲得极其清楚,因为我认为有些人并不了解,而另一些人虽了解却容易遗忘:要全面、一致地控制模型的行为,难度极大。你不能简单地伸手进去说:‘哦,我希望模型少一点道歉。’你当然可以这么做,例如在训练数据中加入‘模型应减少道歉’的指令。但这样一来,在其他某些情境下,模型反而可能变得极其粗鲁,或表现出误导性的过度自信。

英文原文

The second caveat, and I just want to say this super clearly because I think some people don’t know it, others know it, but forget it. It is very difficult to control across the board how the models behave. You cannot just reach in there and say, “Oh, I want the model to apologize less.” You can do that. You can include training data that says, “Oh, the model should apologize less.” But then in some other situation, they end up being super rude or overconfident in a way that’s misleading people.

Dario Amodei48:30

因此,这其中充满了各种权衡取舍。再举一例:曾有一段时期,包括我们自己的模型以及我认为其他公司的模型在内,都过于啰嗦——它们会重复自己,会说得太多。你可以通过惩罚模型‘说话过长’来降低其啰嗦程度。但若以一种粗糙的方式实施该惩罚,就会出现如下情况:当模型编写代码时,有时它会说:‘其余代码写在这里。’

英文原文

So there are all these trade-offs. For example, another thing is if there was a period during which models, ours and I think others as well, were too verbose, they would repeat themselves, they would say too much. You can cut down on the verbosity by penalizing the models for just talking for too long. What happens when you do that, if you do it in a crude way, is when the models are coding, sometimes they’ll say, “Rest of the code goes here,” right?

Dario Amodei48:58

因为它已学会这种‘节省篇幅’的方式,并且在训练数据中见过类似表达。于是这就导致模型在编程时变得所谓‘懒惰’,即只是敷衍道:‘啊,剩下的部分你自己补全吧。’这并非因为我们想节省算力,也不是因为模型在寒假期间变懒了,更不是其他任何曾流传过的阴谋论所致。事实上,仅仅是因为控制模型行为、在所有情境下同时引导其行为,本身就极其困难。

英文原文

Because they’ve learned that that’s the way to economize and that they see it. And then so that leads the model to be so-called lazy in coding where they’re just like, “Ah, you can finish the rest of it.” It’s not because we want to save on compute or because the models are lazy during winter break or any of the other conspiracy theories that have come up. Actually, it’s just very hard to control the behavior of the model, to steer the behavior of the model in all circumstances at once.

Dario Amodei49:28

这就像打地鼠游戏:你按下这边,那边又冒出来,甚至有些变化你根本注意不到、也无法测量。因此,我之所以如此重视未来AI系统的‘宏大对齐’(grand alignment),正是因为这些系统实际上相当不可预测,也的确极难引导和控制。而我们今天所见到的‘改善某一方面,却损害另一方面’的现象,我认为正是未来AI系统控制难题在当下的一种映射,是我们如今便可以着手研究的课题。

英文原文

There’s this whack- a-mole aspect where you push on one thing and these other things start to move as well that you may not even notice or measure. And so one of the reasons that I care so much about grand alignment of these AI systems in the future is actually, these systems are actually quite unpredictable. They’re actually quite hard to steer and control. And this version we’re seeing today of you make one thing better, it makes another thing worse, I think that’s like a present day analog of future control problems in AI systems that we can start to study today.

Dario Amodei50:12

我认为,这种难以精准引导模型行为的困境——即当我们试图将AI系统朝某一方向推动时,它却在我们未预料、也不愿见的其他方向上自行偏移——正是未来趋势的早期征兆。如果我们能妥善解决这一难题,例如:当你要求模型制造并分发天花病毒时,它坚决拒绝;但当你正在攻读病毒学研究生课程时,它又乐于提供帮助——我们该如何同时实现这两种效果?这很难。

英文原文

I think that difficulty in steering the behavior and making sure that if we push an AI system in one direction, it doesn’t push it in another direction in some other ways that we didn’t want. I think that’s an early sign of things to come, and if we can do a good job of solving this problem of you ask the model to make and distribute smallpox and it says no, but it’s willing to help you in your graduate level virology class, how do we get both of those things at once? It’s hard.

Dario Amodei50:48

非常容易走向两个极端中的任意一边,而这是一个多维问题。因此,我认为塑造模型‘人格’这类问题极其困难。我认为我们在这些问题上尚未做到尽善尽美;事实上,我们已做得比所有其他AI公司都更好,但距离完美依然遥不可及。

英文原文

It’s very easy to go to one side or the other and it’s a multidimensional problem. And so I think these questions of shaping the model’s personality, I think they’re very hard. I think we haven’t done perfectly on them. I think we’ve actually done the best of all the AI companies, but still so far from perfect.

Dario Amodei51:08

而且我认为,如果我们能在当前这一高度可控的环境中,成功管控好‘假阳性’与‘假阴性’问题,那么未来当我们真正担忧‘模型是否会具备超强自主性’‘是否能制造极度危险之物’‘是否能自主创建整家公司,且这些公司是否对齐人类价值观’时,我们将更有能力应对。因此,我把当前这项任务既视为棘手挑战,也视作面向未来的良好演练。

英文原文

And I think if we can get this right, if we can control the false positives and false negatives in this very controlled present day environment, we’ll be much better at doing it for the future when our worry is: will the models be super autonomous? Will they be able to make very dangerous things? Will they be able to autonomously build whole companies and are those companies aligned? So I think of this present task as both vexing but also good practice for the future.

Lex Fridman51:40

目前收集用户反馈的最佳方式是什么?不是零散的轶事性数据,而是大规模、系统性的数据,涵盖痛点、亦或其反面——即亮点、积极体验等?是依靠内部测试?特定小组测试?还是A/B测试?哪种方式最有效?

英文原文

What’s the current best way of gathering user feedback? Not anecdotal data, but just large-scale data about pain points or the opposite of pain points, positive things, so on? Is it internal testing? Is it a specific group testing, A/B testing? What works?

Dario Amodei51:59

通常,我们会组织内部‘模型压力测试’(model bashings):整个Anthropic公司——目前员工近一千人——全员参与,竭力尝试‘击垮’模型,以各种方式与之交互。我们还拥有一套评估体系,用于检测‘模型是否出现了此前未曾出现过的拒绝行为’。我们甚至曾专门设计一项‘当然’(certainly)评估,原因在于:某段时间里,模型出现了这样一个恼人的语言习惯——面对范围极广的问题,它总以‘当然,我可以帮您’‘当然,我很乐意为您效劳’‘当然,这是正确的’作为回应。

英文原文

So typically, we’ll have internal model bashings where all of Anthropic… Anthropic is almost 1,000 people. People just try and break the model. They try and interact with it various ways. We have a suite of evals for, “Oh, is the model refusing in ways that it couldn’t?” I think we even had a “certainly” eval because again, at one point, the model had this problem where it had this annoying tick where it would respond to a wide range of questions by saying, “Certainly, I can help you with that. Certainly, I would be happy to do that. Certainly, this is correct.”

Dario Amodei52:34

因此我们设立了‘当然’评估,即统计模型说出‘当然’一词的频率。但请注意,这本质上仍是打地鼠游戏:倘若它把‘当然’换成‘肯定’(definitely)呢?所以,每次我们新增一项评估指标,同时仍需持续运行所有旧有评估项;目前我们已有数百项此类评估。但我们发现,没有任何方法能替代真人与模型的实际交互。

英文原文

And so we had a “certainly” eval, which is: how often does the model say certainly? But look, this is just a whack-a-mole. What if it switches from “certainly” to “definitely”? So every time we add a new eval and we’re always evaluating for all the old things, we have hundreds of these evaluations, but we find that there’s no substitute for a human interacting with it.

Dario Amodei52:56

因此,整个过程非常类似于常规的产品开发流程:我们让Anthropic内部数百名员工对模型进行压力测试;随后开展外部A/B测试;有时还会聘请外包人员进行测试——我们向这些外包人员付费,请他们与模型交互。将所有这些方式综合起来,结果依然不够完美:你仍会观察到一些并不希望看到的行为;你仍会看到模型拒绝一些明显毫无理由拒绝的请求。

英文原文

And so it’s very much like the ordinary product development process. We have hundreds of people within Anthropic bash the model. Then we do external A/B tests. Sometimes we’ll run tests with contractors. We pay contractors to interact with the model. So you put all of these things together and it’s still not perfect. You still see behaviors that you don’t quite want to see. You still see the model refusing things that it just doesn’t make sense to refuse.

Dario Amodei53:25

但我认为,努力应对这一挑战——即阻止模型做出所有人都公认其绝不可为的真正恶劣行为,例如谈论儿童虐待材料之类的内容——这一点上大家意见一致;所有人都认同模型绝不该做这种事。但与此同时,模型也不应以那些愚蠢又笨拙的方式加以拒绝。

英文原文

But I think trying to solve this challenge, trying to stop the model from doing genuinely bad things that everyone agrees it shouldn’t do, everyone agrees that the model shouldn’t talk about, I don’t know, child abuse material. Everyone agrees the model shouldn’t do that, but at the same time, that it doesn’t refuse in these dumb and stupid ways.

Dario Amodei53:49

我认为,尽可能精细地、趋近完美地划定这条界限,本身仍是一项挑战;我们每天都在不断进步,但仍有许多问题亟待解决。再次强调,我愿将此视为未来一项重要挑战的征兆——即如何引导能力更强大的模型。

英文原文

I think drawing that line as finely as possible, approaching perfectly, is still a challenge and we’re getting better at it every day, but there’s a lot to be solved. And again, I would point to that as an indicator of a challenge ahead in terms of steering much more powerful models.

Lex Fridman54:06

您觉得Claude 4.0有朝一日会发布吗?

英文原文

Do you think Claude 4.0 is ever coming out?

Dario Amodei54:11

我不愿对任何命名方案作出承诺,因为倘若我在此宣称‘我们明年将推出Claude 4’,而后我们却决定推倒重来——比如因出现一种新型模型而需另起炉灶——那我就等于自缚手脚。按常规业务节奏,Claude 4理应紧随Claude 3.5之后发布,但在这个变幻莫测的领域里,谁也说不准。

英文原文

I don’t want to commit to any naming scheme because if I say here, “We’re going to have Claude 4 next year,” and then we decide that we should start over because there’s a new type of model, I don’t want to commit to it. I would expect in a normal course of business that Claude 4 would come after Claude 3. 5, but you never know in this wacky field.

Lex Fridman54:34

但这种‘扩展规模’(scaling)的理念仍在持续。

英文原文

But this idea of scaling is continuing.

Dario Amodei54:38

扩展规模仍在持续。我们必将推出比现有模型更强大的新模型,这一点确凿无疑;倘若未能做到,那便意味着我们公司已遭遇严重失败。

英文原文

Scaling is continuing. There will definitely be more powerful models coming from us than the models that exist today. That is certain. Or if there aren’t, we’ve deeply failed as a company.

Lex Fridman54:49

好的。您能解释一下‘负责任扩展政策’(Responsible Scaling Policy)以及人工智能安全等级标准(ASL等级)吗?

英文原文

Okay. Can you explain the responsible scaling policy and the AI safety level standards, ASL levels?

Dario Amodei54:55

尽管我对这些模型带来的益处满怀期待——若谈及《慈爱之机》(Machines of Loving Grace),我们还可就此展开讨论——但我始终担忧其潜在风险,并持续保持警惕。任何人都不应误以为《慈爱之机》一文代表我已不再担忧这些模型的风险;我认为二者实为同一枚硬币的两面。

英文原文

As much as I am excited about the benefits of these models, and we’ll talk about that if we talk about Machines of Loving Grace, I’m worried about the risks and I continue to be worried about the risks. No one should think that Machines of Loving Grace was me saying I’m no longer worried about the risks of these models. I think they’re two sides of the same coin.

Dario Amodei55:16

模型所具备的强大能力,及其在生物学、神经科学、经济发展、治理与和平、乃至经济诸多领域中解决各类问题的潜力,本身亦伴随着风险,对吧?能力越大,责任越重。二者密不可分:强大之物既能行善,亦可作恶。我将这些风险大致划分为若干类别,其中我尤为关注的或许有两个最主要的风险——这并非否认当下已存在其他重要风险,而是当我思考那些可能在最宏大规模上发生之事时,首要考虑的便是我所称的‘灾难性滥用’(catastrophic misuse)。

英文原文

The power of the models and their ability to solve all these problems in biology, neuroscience, economic development, governance and peace, large parts of the economy, those come with risks as well, right? With great power comes great responsibility. The two are paired. Things that are powerful can do good things and they can do bad things. I think of those risks as being in several different categories, perhaps the two biggest risks that I think about. And that’s not to say that there aren’t risks today that are important, but when I think of really the things that would happen on the grandest scale, one is what I call catastrophic misuse.

Dario Amodei55:59

这类滥用涉及网络、生物、放射性、核等领域,一旦严重失控,可能导致成千上万人甚至数百万人伤亡。防范此类风险乃头等要务。在此,我仅作一个简单观察:若审视当今世上那些真正作恶之人,我认为人类之所以迄今尚得庇护,恰恰在于‘极其聪明且受过良好教育者’与‘意图实施极端恐怖行径者’这两类人群之间的交集总体而言非常小。

英文原文

These are misuse of the models in domains like cyber, bio, radiological, nuclear, things that could harm or even kill thousands, even millions of people if they really, really go wrong. These are the number one priority to prevent. And here I would just make a simple observation, which is that the models, if I look today at people who have done really bad things in the world, I think actually humanity has been protected by the fact that the overlap between really smart, well-educated people and people who want to do really horrific things has generally been small.

Dario Amodei56:44

假设某人拥有该领域的博士学位,且拥有一份高薪工作,那么他将失去的东西实在太多。即便假定此人彻底邪恶(而绝大多数人并非如此),他又为何甘愿拿自己的生命、毕生声誉与历史遗产去冒险,以实施真正、真正邪恶之事?倘若此类人数量激增,世界将变得危险得多。因此,我的忧虑在于:AI作为远更智能的主体,或将打破这一相关性。

英文原文

Let’s say I’m someone who I have a PhD in this field, I have a well-paying job. There’s so much to lose. Even assuming I’m completely evil, which most people are not, why would such a person risk their life, risk their legacy, their reputation to do something truly, truly evil? If we had a lot more people like that, the world would be a much more dangerous place. And so my worry is that by being a much more intelligent agent, AI could break that correlation.

Dario Amodei57:21

因此,我对此确有严重忧虑。我相信这些忧虑可以被防范。但作为对《慈爱之机》一文的补充说明,我想强调:风险依然严峻。第二类风险则是‘自主性风险’(autonomy risks),即模型可能在无人干预的情况下自行行动——尤其当我们赋予其较以往更强的自主权,例如让其承担更广泛的任务,如编写整套代码库,乃至未来某日实际上运营整家公司时,其行动自由度已足够大——它们是否仍在切实执行我们真正期望之事?

英文原文

And so I do have serious worries about that. I believe we can prevent those worries. But I think as a counterpoint to Machines of Loving Grace, I want to say that there’s still serious risks. And the second range of risks would be the autonomy risks, which is the idea that models might, on their own, particularly as we give them more agency than they’ve had in the past, particularly as we give them supervision over wider tasks like writing whole code bases or someday even effectively operating entire companies, they’re on a long enough leash. Are they doing what we really want them to do?

Dario Amodei58:00

我们甚至难以详尽理解其具体行为,更遑论加以控制。正如我此前所言,早期迹象已表明:精确划定模型‘应做之事’与‘不应做之事’的边界极为困难;若偏向一侧,结果便令人厌烦且毫无用处;若偏向另一侧,则引发其他不良行为;修复某一问题,往往又催生新的问题。

英文原文

It’s very difficult to even understand in detail what they’re doing, let alone control it. And like I said, these early signs that it’s hard to perfectly draw the boundary between things the model should do and things the model shouldn’t do that if you go to one side, you get things that are annoying and useless and you go to the other side, you get other behaviors. If you fix one thing, it creates other problems.

Dario Amodei58:25

我们正日益精进于解决此类问题。我认为这并非无解之题。它更像一门科学,如同飞机安全、汽车安全或药品安全那样。我不认为我们遗漏了任何重大环节,只是需要不断提升对这些模型的管控能力。以上便是我所担忧的两类风险。至于我们的‘负责任扩展计划’(RSP),我承认,这已是对你提问的一个相当冗长的回答。

英文原文

We’re getting better and better at solving this. I don’t think this is an unsolvable problem. I think this is a science like the safety of airplanes or the safety of cars or the safety of drugs. I don’t think there’s any big thing we’re missing. I just think we need to get better at controlling these models. And so these are the two risks I’m worried about. And our responsible scaling plan, which I’ll recognize is a very long-winded answer to your question.

Lex Fridman58:49

我喜欢!我非常喜欢!

英文原文

I love it. I love it.

Dario Amodei58:51

我们的‘负责任扩展计划’正是为应对上述两类风险而设计。因此,每次我们开发一款新模型,本质上都会测试其实施这两类有害行为的能力。若稍作回溯,我认为AI系统面临一个有趣的困境:它们目前尚不足以引发此类灾难。我不知道它们未来是否会引发此类灾难,这种可能性或许并不存在。

英文原文

Our responsible scaling plan is designed to address these two types of risks. And so every time we develop a new model, we basically test it for its ability to do both of these bad things. So if I were to back up a little bit, I think we have an interesting dilemma with AI systems where they’re not yet powerful enough to present these catastrophes. I don’t know if they’ll ever present these catastrophes. It’s possible they won’t.

Dario Amodei59:22

但值得忧虑的理由、风险存在的依据已足够充分,足以促使我们立即行动;况且,模型能力正以极快速度提升。约一年前,我在参议院作证时曾指出,我们可能在两到三年内面临严重的生物风险。此后进展一如预期。因此,我们面临一种奇特状况:这些风险虽尚未显现、尚不存在,却如幽灵般迫近——只因模型进化速度实在太快。

英文原文

But the case for worry, the case for risk is strong enough that we should act now and they’re getting better very, very fast. I testified in the Senate that we might have serious bio risks within two to three years. That was about a year ago. Things have proceeded apace. So we have this thing where it’s surprisingly hard to address these risks because they’re not here today, they don’t exist. They’re like ghosts, but they’re coming at us so fast because the models are improving so fast.

Dario Amodei59:56

那么,该如何应对一种‘当下尚不存在、却正以极快速度向我们逼近’的事物?为此,我们联合METR等组织及保罗·克里斯蒂亚诺(Paul Christiano)等人,提出了一种解决方案:我们需要能预示风险临近的测试手段,即一套早期预警系统。因此,每次我们推出新模型,均会测试其执行CBRN(化学、生物、放射性、核)相关任务的能力,同时测试其独立自主完成任务的能力。

英文原文

So how do you deal with something that’s not here today, doesn’t exist, but is coming at us very fast? So the solution we came up with for that, in collaboration with people like the organization METR and Paul Christiano is what you need for that are you need tests to tell you when the risk is getting close. You need an early warning system. And so every time we have a new model, we test it for its capability to do these CBRN tasks as well as testing it for how capable it is of doing tasks autonomously on its own.

Dario Amodei1:00:35

而在我们最近一两个月发布的最新版RSP中,我们测试自主性风险的方式,是评估AI模型开展AI自身研究的能力;一旦AI模型能够从事AI研究,它们便真正、真正实现了自主。这一临界点在诸多其他方面亦具有重要意义。那么,针对这些测试任务,我们究竟采取何种措施?RSP基本构建了一种‘如果—那么’(if-then)结构:即若模型通过某项特定能力门槛,则须对其施加相应的一系列安全与安保要求。

英文原文

And in the latest version of our RSP, which we released in the last month or two, the way we test autonomy risks is the AI model’s ability to do aspects of AI research itself, which when the AI models can do AI research, they become truly, truly autonomous. And that threshold is important for a bunch of other ways. And so what do we then do with these tasks? The RSP basically develops what we’ve called an if-then structure, which is if the models pass a certain capability, then we impose a certain set of safety and security requirements on them.

Dario Amodei1:01:16

当前模型属于所谓‘ASL-2’级别。ASL-1则适用于明显不构成任何自主性或滥用风险的系统。例如,国际象棋程序‘深蓝’(Deep Blue)即属ASL-1:显而易见,它只能用于下棋——它本就专为下棋而设计,没人会用它发动高超的网络攻击,更不会让它失控并接管世界。

英文原文

So today’s models are what’s called ASL-2. Models that were ASL-1 is for systems that manifestly don’t pose any risk of autonomy or misuse. So for example, a chess playing bot, Deep Blue would be ASL-1. It’s just manifestly the case that you can’t use Deep Blue for anything other than chess. It was just designed for chess. No one’s going to use it to conduct a masterful cyber attack or to run wild and take over the world.

Dario Amodei1:01:47

ASL-2对应当前AI系统:我们已对其能力进行测量,认为这些系统尚不足以自主自我复制,或独立执行大量任务;亦不足以提供超出谷歌搜索所能获取范围的、关于CBRN风险及CBRN武器制造的实质性信息。事实上,它们有时确实能提供超越搜索引擎的信息,但这些信息无法被整合串联,亦无法形成端到端的、足以构成危险的完整链条。

英文原文

ASL-2 is today’s AI systems where we’ve measured them and we think these systems are simply not smart enough to autonomously self-replicate or conduct a bunch of tasks and also not smart enough to provide meaningful information about CBRN risks and how to build CBRN weapons above and beyond what can be known from looking at Google. In fact, sometimes they do provide information above and beyond a search engine, but not in a way that can be stitched together, not in a way that end-to-end is dangerous enough.

Dario Amodei1:02:26

因此,ASL-3将标志着模型能力达到一个临界点:其辅助能力足以增强非国家行为体(non-state actors)的实力。不幸的是,国家行为体目前已能在相当高的熟练度上实施诸多极度危险且具破坏性的行为;关键区别在于,非国家行为体尚不具备此种能力。因此,一旦迈入ASL-3阶段,我们将采取特殊安保措施,确保足以防止模型遭非国家行为体窃取及部署后被滥用;我们还将必须配备针对这些特定领域的强化过滤机制。

英文原文

So ASL-3 is going to be the point at which the models are helpful enough to enhance the capabilities of non-state actors, right? State actors can already do, unfortunately, to a high level of proficiency, a lot of these very dangerous and destructive things. The difference is that non-state actors are not capable of it. And so when we get to ASL-3, we’ll take special security precautions designed to be sufficient to prevent theft of the model by non-state actors and misuse of the model as it’s deployed. We’ll have to have enhanced filters targeted at these particular areas.

Lex Fridman1:03:07

网络、生物、核。

英文原文

Cyber, bio, nuclear.

Dario Amodei1:03:09

网络、生物、核领域以及模型自主性——后者并非主要涉及误用风险,而更多是指模型自身可能主动做出有害行为的风险。ASL-4 指的是模型能力发展到足以增强已有专业知识的国家行为体之能力,或本身即成为此类风险的主要来源。倘若你有意引发此类风险,那么借助模型将是主要途径。此外,在自主性维度上,ASL-4 还意味着借助 AI 模型在 AI 研究能力方面实现一定程度的加速。

英文原文

Cyber, bio, nuclear and model autonomy, which is less a misuse risk and more a risk of the model doing bad things itself. ASL-4, getting to the point where these models could enhance the capability of a already knowledgeable state actor and/or become the main source of such a risk. If you wanted to engage in such a risk, the main way you would do it is through a model. And then I think ASL-4 on the autonomy side, it’s some amount of acceleration in AI research capabilities with an AI model.

Dario Amodei1:03:45

而 ASL-5 则指模型真正具备了卓越能力,其在执行上述任一任务方面均能超越人类。因此,这种‘如果—那么’结构的承诺,其核心意图在于表明:‘看,我并不确定——我已与这些模型共事多年,也已担忧相关风险多年。事实上,滥发警报十分危险;事实上,断言某模型存在风险也十分危险——人们审视该模型后会说:“这显然不具危险性。”’再次强调,风险的微妙性不在于当下已然显现,而在于它正以极快速度向我们逼近。

英文原文

And then ASL-5 is where we would get to the models that are truly capable that it could exceed humanity in their ability to do any of these tasks. And so the point of the if-then structure commitment is basically to say, “Look, I don’t know, I’ve been working with these models for many years and I’ve been worried about risk for many years. It’s actually dangerous to cry wolf. It’s actually dangerous to say this model is risky. And people look at it and they say this is manifestly not dangerous.” Again, it’s the delicacy of the risk isn’t here today, but it’s coming at us fast.

Dario Amodei1:04:27

你该如何应对这种情况?这对风险规划者而言确实极为棘手。因此,这种‘如果—那么’结构本质上是在表明:‘看,我们不想激怒一大群人,也不想因对当前尚无危险的模型施加过重负担,而损害自身参与相关讨论的话语权。’所以,这种‘如果—那么’式的触发承诺,本质上正是为应对此类困境而设——即仅当确凿证明模型具有危险性时,才对其实施严格管控。

英文原文

How do you deal with that? It’s really vexing to a risk planner to deal with it. And so this if-then structure basically says, “Look, we don’t want to antagonize a bunch of people, we don’t want to harm our own ability to have a place in the conversation by imposing these very onerous burdens on models that are not dangerous today.” So the if-then, the trigger commitment is basically a way to deal with this. It says you clamp down hard when you can show the model is dangerous.

Dario Amodei1:04:58

当然,与此配套的必须是足够宽裕的缓冲阈值,以避免高概率地漏判危险。这一框架并非完美无缺,我们已不得不对其进行调整。就在几周前,我们刚推出一个新版本;未来,我们甚至可能每年多次发布更新版本,因为从技术层面和组织研究视角来看,制定出恰如其分的政策实属不易。但这就是我们的提案:采用‘如果—那么’式承诺与触发机制,旨在当前最大限度减少负担与误报,同时确保当危险真正来临时,能作出恰当响应。

英文原文

And of course, what has to come with that is enough of a buffer threshold that you’re not at high risk of missing the danger. It’s not a perfect framework. We’ve had to change it. We came out with a new one just a few weeks ago and probably going forward, we might release new ones multiple times a year because it’s hard to get these policies right technically, organizationally from a research perspective. But that is the proposal, if-then commitments and triggers in order to minimize burdens and false alarms now, but really react appropriately when the dangers are here.

Lex Fridman1:05:37

您认为 ASL-3(即多项触发条件被激活)的时间线是怎样的?ASL-4 的时间线又如何?

英文原文

What do you think the timeline for ASL-3 is where several of the triggers are fired? And what do you think the timeline is for ASL-4?

Dario Amodei1:05:44

是的,这在公司内部正激烈争论。我们正积极筹备 ASL-3 的安全措施及部署措施。我暂不详述细节,但我们在两方面均已取得大量进展,并且我认为我们很快就能做好准备。若明年即达到 ASL-3,我完全不会感到意外;甚至有人担心今年就可能达到,这种可能性依然存在。虽然很难断言,但我若听说要等到 2030 年,那将令我极度震惊——我认为实际时间点会远早于此。

英文原文

Yeah. So that is hotly debated within the company. We are working actively to prepare ASL-3 security measures as well as ASL-3 deployment measures. I’m not going to go into detail, but we’ve made a lot of progress on both and we’re prepared to be, I think, ready quite soon. I would not be surprised at all if we hit ASL-3 next year. There was some concern that we might even hit it this year. That’s still possible. That could still happen. It’s very hard to say, but I would be very, very surprised if it was 2030. I think it’s much sooner than that.

Lex Fridman1:06:24

因此,既有用于检测的协议(即‘如果—那么’机制),也有针对检测结果的响应协议。

英文原文

So there’s protocols for detecting it, the if-then and then there’s protocols for how to respond to it.

Dario Amodei1:06:31

是的。

英文原文

Yes.

Lex Fridman1:06:32

那么,后者——即响应协议——难度如何?

英文原文

How difficult is the second, the latter?

Dario Amodei1:06:34

是的,我认为对于 ASL-3 而言,重点主要在于安全性,以及模型部署时针对极少数特定领域的过滤机制。因为在 ASL-3 阶段,模型尚不具备自主性,因此即使模型在内部部署,也不必担忧其自身会做出有害行为。所以,我认为 ASL-3 措施虽不能说是简单直接,但确实严谨,且更易于进行逻辑推演与分析。

英文原文

Yeah. I think for ASL-3, it’s primarily about security and about filters on the model relating to a very narrow set of areas when we deploy the model. Because at ASL-3, the model isn’t autonomous yet. And so you don’t have to worry about the model itself behaving in a bad way even when it’s deployed internally. So I think the ASL- 3 measures are, I won’t say straightforward, they’re rigorous, but they’re easier to reason about.

Dario Amodei1:07:06

我认为一旦进入 ASL-4 阶段,我们便开始担忧模型是否已足够聪明,以至于可能在测试中故意藏拙、隐瞒真实能力,甚至可能在测试中说谎。此前已有相关成果发布,涉及‘休眠特工’现象;近期还有一篇论文探讨:‘模型能否误导外界对其能力的评估,刻意伪装成比实际能力更弱的状态?’因此,我认为在 ASL-4 阶段,一个重要组成部分将是采用除直接与模型交互之外的其他手段。

英文原文

I think once we get to ASL-4, we start to have worries about the models being smart enough that they might sandbag tests, they might not tell the truth about tests. We had some results came out about sleeper agents and there was a more recent paper about, “Can the models mislead attempts to sandbag their own abilities, present themselves as being less capable than they are?” And so I think with ASL-4, there’s going to be an important component of using other things than just interacting with the models.

Dario Amodei1:07:43

例如,可借助可解释性技术,或探查模型内部隐藏的思维链路——即需深入模型内部,通过某种不易被模型自身操控的其他机制加以验证,从而确认模型是否确实具备某项特性。目前我们仍在推进 ASL-4 相关工作。RSP(负责任扩展政策)的一项特性在于:我们不会预先定义 ASL-4,而是待达成 ASL-3 后再行制定。我们认为这一决策已被证明十分明智,因为即便对于 ASL-3,其具体细节仍难以精确把握;我们希望尽可能争取充足时间,务求将各项措施做到万无一失。

英文原文

For example, interpretability or hidden chains of thought where you have to look inside the model and verify via some other mechanism that is not as easily corrupted as what the model says, that the model indeed has some property. So we’re still working on ASL-4. One of the properties of the RSP is that we don’t specify ASL-4 until we’ve hit ASL-3. And I think that’s proven to be a wise decision because even with ASL-3, again, it’s hard to know this stuff in detail, and we want to take as much time as we can possibly take to get these things right.

Lex Fridman1:08:23

因此,在 ASL-3 阶段,作恶者将是人类。

英文原文

So for ASL-3, the bad actor will be the humans.

Dario Amodei1:08:26

人类,没错。

英文原文

Humans, yes.

Lex Fridman1:08:27

因此,这方面会略显……

英文原文

And so there’s a little bit more…

Dario Amodei1:08:29

而在 ASL-4 阶段,作恶者则既包括人类,也包括模型本身,我认为是这样。

英文原文

For ASL- 4, it’s both, I think.

Lex Fridman1:08:31

是两者兼具。因此,欺骗行为将成为关键问题,而此时机制性可解释性(mechanistic interpretability)便派上用场;我们亦希望用于该目的的技术不会被模型所获取。

英文原文

It’s both. And so deception, and that’s where mechanistic interpretability comes into play, and hopefully the techniques used for that are not made accessible to the model.

Dario Amodei1:08:42

是的。当然,你完全可以将机制性可解释性工具直接接入模型本身,但如此一来,它便不再能作为模型状态的可靠指示器。此外,还有不少异乎寻常的情形可能导致其不可靠——例如,若模型足够聪明,竟能跨设备跳转并读取你正在检查其内部状态的代码。我们已考虑过其中一些情形,但认为它们过于离奇;不过,我们确实存在若干方法可使其发生概率大幅降低。但总体而言,你希望将机制性可解释性保留为一套独立于模型训练过程之外的验证集或测试集。

英文原文

Yeah. Of course, you can hook up the mechanistic interpretability to the model itself, but then you’ve lost it as a reliable indicator of the model state. There are a bunch of exotic ways you can think of that it might also not be reliable, like if the model gets smart enough that it can jump computers and read the code where you’re looking at its internal state. We’ve thought about some of those. I think they’re exotic enough. There are ways to render them unlikely. But yeah, generally, you want to preserve mechanistic interpretability as a verification set or test set that’s separate from the training process of the model.

Lex Fridman1:09:19

您看,随着这些模型对话能力日益精进、智能水平持续提升,社会工程学攻击也将成为一种威胁,因为它们可能对各公司内部的工程师极具说服力。

英文原文

See, I think as these models become better and better conversation and become smarter, social engineer becomes a threat too because they could start being very convincing to the engineers inside companies.

Dario Amodei1:09:30

哦,是的,是的。我们在生活中已目睹过大量人类煽动者的实例,而人们也担忧模型同样可能施展此类手段。

英文原文

Oh, yeah. Yeah. We’ve seen lots of examples of demagoguery in our life from humans, and there’s a concern that models could do that as well.

Lex Fridman1:09:40

Claude 日益强大的一种体现,是它如今已能执行某些具身智能(agentic)任务,例如计算机操作。此外,在 claude.ai 自身的沙盒环境中也存在相应分析。但让我们聚焦于计算机操作——这在我看来极其令人振奋:你只需向 Claude 下达一项任务,它便能采取一系列操作、自行理清思路,并获得对你计算机的访问权限……

英文原文

One of the ways that Claude has been getting more and more powerful is it’s now able to do some agentic stuff, computer use. There’s also an analysis within the sandbox of Claude.ai itself. But let’s talk about computer use. That seems to me super exciting that you can just give Claude a task and it takes a bunch of actions, figures it out, and has access to the…

Lex Fridman1:10:00

……采取一系列操作、理清思路,并通过截图方式访问你的计算机。那么,能否请您解释一下其运作原理,以及未来发展方向?

英文原文

… a bunch of actions, figures it out and has access to your computer through screenshots. So can you explain how that works and where that’s headed?

Dario Amodei1:10:10

是的,其实原理相当简单。自今年三月 Claude 3 发布以来,Claude 长期具备分析图像并以文本形式回应的能力。我们新增的唯一功能,是允许这些图像为计算机屏幕截图;相应地,我们对模型进行了训练,使其能输出屏幕上可供点击的位置坐标,和/或键盘上可供按下的按键指令,从而执行具体操作。事实证明,仅需额外投入相对有限的训练资源,模型即可在此任务上表现得相当出色。这正是泛化能力的一个绝佳例证。人们有时会说:‘只要抵达近地轨道,你就已走完通往任意地点旅程的一半’,因为挣脱地球引力势阱所需能量巨大。同理,若拥有一个强大预训练模型,我感觉你在智能空间中也已走完一半路程。因此,让 Claude 掌握此项能力实际并未耗费太多额外资源。你只需将其置于循环之中:向模型提供一张截图,告知其应点击何处;再提供下一张截图,再告知其应点击何处——如此往复,便形成了一种近乎三维视频交互的完整模式,使模型得以完成所有此类任务。我们展示的演示案例中,它能填写电子表格、与网站互动、启动各类程序,且兼容不同操作系统——Windows、Linux、Mac 均可。因此,我认为这一切都极为激动人心。我必须指出:理论上,通过直接向模型提供驱动计算机屏幕的 API 接口,你本也能实现相同功能;但当前方案真正大幅降低了使用门槛——许多用户要么无法接触这些 API,要么调用它们耗时甚久。

英文原文

Yeah. It’s actually relatively simple. So Claude has had for a long time, since Claude 3 back in March, the ability to analyze images and respond to them with text. The only new thing we added is those images can be screenshots of a computer and in response, we train the model to give a location on the screen where you can click and/or buttons on the keyboard, you can press in order to take action. And it turns out that with actually not all that much additional training, the models can get quite good at that task. It’s a good example of generalization. People sometimes say if you get to lower earth orbit, you’re halfway to anywhere because of how much it takes to escape the gravity well. If you have a strong pre-trained model, I feel like you’re halfway to anywhere in terms of the intelligence space. And so actually, it didn’t take all that much to get Claude to do this. And you can just set that in a loop, give the model a screenshot, tell it what to click on, give it the next screenshot, tell it what to click on and that turns into a full kind of almost 3D video interaction of the model and it’s able to do all of these tasks. We showed these demos where it’s able to fill out spreadsheets, it’s able to kind of interact with a website, it’s able to open all kinds of programs, different operating systems, Windows, Linux, Mac. So I think all of that is very exciting. I will say, while in theory there’s nothing you could do there that you couldn’t have done through just giving the model the API to drive the computer screen, this really lowers the barrier. And there’s a lot of folks who either aren’t in a position to interact with those APIs or it takes them a long time to do.

Dario Amodei1:12:00

屏幕本身只是一种通用界面,交互起来要容易得多。因此,我预计随着时间推移,这将降低大量门槛。坦白讲,当前模型仍存在诸多不足,我们在博客中也坦诚指出了这一点:它会出错,会误点。我们特意谨慎地提醒用户:‘嘿,你不能把这个东西丢在电脑上连续运行几分钟、几十分钟;你必须为它设定边界和护栏。’我认为,这正是我们首先以API形式发布该模型、而非直接交到消费者手中并赋予其对计算机的控制权的原因之一。但我确实深感,将这些能力尽快推向外界至关重要。随着模型日益强大,我们必须认真思考:如何安全地使用这些能力?如何防止它们被滥用?

英文原文

It’s just the screen is just a universal interface that’s a lot easier to interact with. And so I expect over time, this is going to lower a bunch of barriers. Now, honestly, the current model has, it leaves a lot still to be desired and we were honest about that in the blog. It makes mistakes, it misclicks. We were careful to warn people, “Hey, you can’t just leave this thing to run on your computer for minutes and minutes. You got to give this thing boundaries and guardrails.” And I think that’s one of the reasons we released it first in an API form rather than just hand the consumer and give it control of their computer. But I definitely feel that it’s important to get these capabilities out there. As models get more powerful, we’re going to have to grapple with how do we use these capabilities safely. How do we prevent them from being abused?

Dario Amodei1:12:54

我认为,在能力尚有限时就发布该模型,对于实现上述目标非常有帮助。自发布以来,已有不少客户迅速采用——比如Replit,或许是其中部署速度最快的一家——并以各种方式加以利用。人们已为Windows桌面、Mac及Linux机器搭建了各类演示系统。所以,这确实令人振奋。我想,与任何新技术一样,它既带来了令人兴奋的新能力,也伴随着这些新能力而来的挑战:我们该如何确保模型安全可靠,使其真正按人类意愿行事?这与其他所有技术面临的问题如出一辙,本质是同一场张力。

英文原文

And I think releasing the model while the capabilities are still limited is very helpful in terms of doing that. I think since it’s been released, a number of customers, I think Replit was maybe one of the most quickest to deploy things, have made use of it in various ways. People have hooked up demos for Windows desktops, Macs, Linux machines. So yeah, it’s been very exciting. I think as with anything else, it comes with new exciting abilities and then with those new exciting abilities, we have to think about how to make the model safe, reliable, do what humans want them to do. It’s the same story for everything. Same thing. It’s that same tension.

Lex Fridman1:13:51

但此处潜在用例的可能性之广,实在令人惊叹。那么,为了让它在未来真正表现优异,需要在预训练模型现有能力基础上,额外投入多少精力?是否需开展更多后训练、基于人类反馈的强化学习(RLHF)、监督式微调,或专门针对智能体任务生成合成数据?

英文原文

But the possibility of use cases here, just the range is incredible. So how much to make it work really well in the future? How much do you have to specially kind of go beyond what the pre-trained model is doing, do more post-training, RLHF or supervised fine-tuning or synthetic data just for the agentive stuff?

Dario Amodei1:14:10

是的。从宏观层面讲,我们有意持续大力投入,不断提升模型性能。我们观察到某些基准测试中,此前模型仅能在6%的情况下完成任务,而我们的模型如今可达14%或22%。是的,我们希望最终达到人类水平的可靠性——即80%、90%,就像其他领域一样。我们正沿着与SWE-bench相同的演进曲线前行;我估计,一年后,模型将能极为可靠地完成此类任务。但一切总得从起点开始。

英文原文

Yeah. I think speaking at a high level, it’s our intention to keep investing a lot in making the model better. I think we look at some of the benchmarks where previous models were like, “Oh, could do it 6% of the time,” and now our model would do it 14 or 22% of the time. And yeah, we want to get up to the human level reliability of 80, 90% just like anywhere else. We’re on the same curve that we were on with SWE-bench where I think I would guess a year from now, the models can do this very, very reliably. But you got to start somewhere.

Lex Fridman1:14:41

那么,您认为仅靠目前所采用的方法,就有可能达到人类水平的90%可靠性,还是必须针对计算机使用场景进行专门优化?

英文原文

So you think it’s possible to get to the human level 90% basically doing the same thing you’re doing now or it has to be special for computer use?

Dario Amodei1:14:49

这取决于您所说的‘专门’具体指什么——泛泛而言,我总体上认为,我们当前用于训练模型的各类技术,若像此前在代码、通用模型、图像输入、语音等领域那样加倍投入,同样有望在此处实现规模化扩展。

英文原文

It depends what you mean by special and special in general, but I generally think the same kinds of techniques that we’ve been using to train the current model, I expect that doubling down on those techniques in the same way that we have for code, for models in general, for image input, for voice, I expect those same techniques will scale here as they have everywhere else,

Lex Fridman1:15:18

但此举相当于赋予Claude行动能力,因此既能实现许多极为强大的功能,也可能造成巨大损害。

英文原文

But this is giving the power of action to Claude and so you could do a lot of really powerful things, but you could do a lot of damage also.

Dario Amodei1:15:27

是的,是的。我们对此一直高度警觉。实际上,我的观点是:计算机使用并非像CBRN(化学、生物、放射性、核)或自主性能力那样的根本性新能力;它更类似于为模型打开一扇窗口,使其得以运用并施展既有能力。因此,我们回溯至自身的RSP(负责任缩放政策)框架来思考:该模型本身所执行的任何操作,本质上都不会从RSP角度增加风险;但随着模型日益强大,拥有此项能力可能使风险加剧——一旦模型具备ASL-3或ASL-4级别的认知能力,这项能力或许恰恰成为突破约束、付诸行动的关键因素。因此,未来我们势必会对这种交互模态开展RSP相关测试,并将持续推进此类测试。我认为,在模型尚未具备超强能力之前,先行学习与探索此项能力,或许更为妥当。

英文原文

Yeah, yeah. No and we’ve been very aware of that. Look, my view actually is computer use isn’t a fundamentally new capability like the CBRN or autonomy capabilities are. It’s more like it kind of opens the aperture for the model to use and apply its existing abilities. And so the way we think about it, going back to our RSP, is nothing that this model is doing inherently increases the risk from an RSP perspective, but as the models get more powerful, having this capability may make it scarier once it has the cognitive capability to do something at the ASL-3 and ASL-4 level, this may be the thing that kind of unbounds it from doing so. So going forward, certainly this modality of interaction is something we have tested for and that we will continue to test for an RSP going forward. I think it’s probably better to learn and explore this capability before the model is super capable

Lex Fridman1:16:33

是的。还有许多有趣的攻击方式,例如提示注入(prompt injection),因为如今您已拓宽了攻击面,可通过屏幕上显示的内容实施提示注入。倘若该能力变得越来越实用,那么向模型注入恶意内容的动机与收益也将随之增长。例如,若模型访问某网页,注入内容可能是无害的广告,但也可能是有害内容,对吧?

英文原文

Yeah. And there’s a lot of interesting attacks like prompt injection because now you’ve widened the aperture so you can prompt inject through stuff on screen. So if this becomes more and more useful, then there’s more and more benefit to inject stuff into the model. If it goes to certain web page, it could be harmless stuff like advertisements or it could be harmful stuff, right?

Dario Amodei1:16:53

是的,我们已深入思考过垃圾信息、验证码、大规模……这里透露一个秘密:若您发明了一项新技术,其最初遭遇的滥用未必是最严重的,但最先出现的滥用行为,往往是诈骗——只是些小打小闹式的诈骗。

英文原文

Yeah, we’ve thought a lot about things like spam, CAPTCHA, mass… One secret, I’ll tell you, if you’ve invented a new technology, not necessarily the biggest misuse, but the first misuse you’ll see, scams, just petty scams.

Lex Fridman1:17:10

是的。

英文原文

Yeah.

Dario Amodei1:17:13

这事儿古已有之——人类彼此诈骗,可谓源远流长。每次出现新技术,都得面对这类问题。

英文原文

It’s like a thing as old, people scamming each other, it’s this thing as old as time. And it’s just every time, you got to deal with it.

Lex Fridman1:17:21

说起来虽近乎可笑,却是事实:随着人工智能整体愈发智能,机器人与垃圾信息泛滥这类问题也愈发严重……

英文原文

It’s almost silly to say, but it’s true, sort of bots and spam in general is a thing as it gets more and more intelligent-

Dario Amodei1:17:29

是的,是的。

英文原文

Yeah, yeah.

Lex Fridman1:17:29

……对抗难度也越来越大。

英文原文

… it’s harder and harder to fight it.

Dario Amodei1:17:32

正如我所说,世上本就存在大量小罪犯,而每项新技术,不过是为他们提供了又一种干蠢事、行恶事的新途径。

英文原文

Like I said, there are a lot of petty criminals in the world and it’s like every new technology is a new way for petty criminals to do something stupid and malicious.

Lex Fridman1:17:45

关于沙盒化(sandboxing)是否有任何构想?沙盒化任务的难度究竟如何?

英文原文

Is there any ideas about sandboxing it? How difficult is the sandboxing task?

Dario Amodei1:17:49

是的,我们在训练阶段即实施沙盒化。例如,训练期间我们并未让模型接入互联网。我认为,训练阶段接入互联网恐怕是个糟糕主意,因为模型可能随时调整自身策略、改变行为方式,进而对现实世界产生影响。至于实际部署模型,则视具体应用场景而定:有时您确实希望模型在现实世界中执行操作。当然,您始终可在外部设置防护措施与护栏,例如声明:‘好吧,该模型不得将我的任何文件从我的计算机或网络服务器转移至任何其他地方。’

英文原文

Yeah, we sandbox during training. So for example, during training we didn’t expose the model to the internet. I think that’s probably a bad idea during training because the model can be changing its policy, it can be changing what it’s doing and it’s having an effect in the real world. In terms of actually deploying the model, it kind of depends on the application. Sometimes you want the model to do something in the real world. But of course, you can always put guard, you can always put guard rails on the outside. You can say, “Okay, well, this model’s not going to move data from my, the model’s not going to move any files from my computer or my web server to anywhere else.”

Dario Amodei1:18:27

再次谈及沙盒化——当我们迈向ASL-4级别时,所有这些预防措施都将失去意义。一旦涉及ASL-4,理论上便存在担忧:模型可能足够聪明,从而突破任何沙盒限制。此时,我们必须转向机制可解释性(mechanistic interpretability)研究。若要构建沙盒,其设计必须具备数学可证明性。这已完全不同于我们当下所应对的模型世界。

英文原文

Now, when you talk about sandboxing, again, when we get to ASL-4, none of these precautions are going to make sense there. When you talk about ASL-4, you’re then, the model is being, there’s theoretical worry the model could be smart enough to kind of break it to out of any box. And so there, we need to think about mechanistic interpretability. If we’re going to have a sandbox, it would need to be a mathematically provable. That’s a whole different world than what we’re dealing with with the models today.

Lex Fridman1:19:01

是的,即构建一个ASL-4级人工智能系统无法逃脱的‘盒子’的科学。

英文原文

Yeah, the science of building a box from which ASL-4 AI system cannot escape.

Dario Amodei1:19:08

我认为这恐怕并非正确路径。相较试图遏制一个目标不一致的模型、防止其逃脱,更优路径应是:从设计之初就构建正确的模型,或建立一种可深入模型内部、验证其属性的闭环机制,从而获得机会反复检验、迭代优化,最终真正达成目标。我认为,管控不良模型,远不如直接构建优质模型来得有效。

英文原文

I think it’s probably not the right approach. I think the right approach, instead of having something unaligned that you’re trying to prevent it from escaping, I think it’s better to just design the model the right way or have a loop where you look inside the model and you’re able to verify properties and that gives you an opportunity to tell, iterate and actually get it right. I think containing bad models is a much worse solution than having good models.

Lex Fridman1:19:36

请允许我谈谈监管问题:监管在保障人工智能安全方面扮演何种角色?例如,能否介绍最终遭州长否决的加州人工智能监管法案SB 1047?该法案总体而言有何利弊?

英文原文

Let me ask about regulation. What’s the role of regulation in keeping AI safe? So for example, can you describe California AI regulation bill SB 1047 that was ultimately vetoed by the governor? What are the pros and cons of this bill in general?

Dario Amodei1:19:50

是的,我们最终向该法案提交了一些建议,其中部分建议被采纳;我们对此法案的整体评价也日趋积极——尽管它仍存在一些缺陷。当然,该法案最终被否决。从宏观层面看,我认为该法案背后的一些核心理念,与我们RSP(负责任缩放政策)的理念颇为相似。我认为,无论加州、联邦政府,抑或其他国家与州,均亟需通过此类监管法规。我可进一步阐述为何我认为此事如此重要。因此,我对我们的RSP持肯定态度——它并非完美无缺,仍需大量迭代完善;但它确已成为一项有力的推动机制,促使公司严肃对待相关风险,将其纳入产品规划流程,并使之真正成为Anthropic全员工作的核心重点;同时确保公司近一千名员工(目前几乎已达千人规模)充分理解:此事乃公司最高优先事项之一,甚至就是最高优先事项。

英文原文

Yes, we ended up making some suggestions to the bill. And then some of those were adopted and we felt, I think, quite positively about the bill by the end of that, it did still have some downsides. And of course, it got vetoed. I think at a high level, I think some of the key ideas behind the bill are I would say similar to ideas behind our RSPs. And I think it’s very important that some jurisdiction, whether it’s California or the federal government and/or other countries and other states, passes some regulation like this. And I can talk through why I think that’s so important. So I feel good about our RSP. It’s not perfect. It needs to be iterated on a lot. But it’s been a good forcing function for getting the company to take these risks seriously, to put them into product planning, to really make them a central part of work at Anthropic and to make sure that all of a thousand people, and it’s almost a thousand people now at Anthropic, understand that this is one of the highest priorities of the company, if not the highest priority.

Dario Amodei1:20:58

但首先,仍有一些公司尚未建立类似RSP(负责任规模扩展)的机制,比如OpenAI;谷歌是在Anthropic推出此类机制数月后才跟进采用的,而其他一些公司则根本未设立这类机制。因此,倘若部分公司采纳了这些机制,而另一些公司没有,就极易形成这样一种局面:某些风险具有特殊性质——即便五家公司中有三家采取了安全措施,只要另外两家不安全,就会产生负面外部性。我认为这种缺乏统一性的状况,对我们这些已投入大量精力、极为审慎地设计并落实相关流程的人来说是不公平的。第二点是,我认为不能仅凭信任,就指望这些公司自发遵守其自愿制定的计划。对吧?我个人倾向于相信Anthropic会恪守承诺,我们竭尽所能去践行;我们的RSP由长期公益信托机构进行监督,因此我们确实在尽一切努力遵守自身制定的RSP。

英文原文

But one, there are still some companies that don’t have RSP like mechanisms, like OpenAI, Google did adopt these mechanisms a couple months after Anthropic did, but there are other companies out there that don’t have these mechanisms at all. And so if some companies adopt these mechanisms and others don’t, it’s really going to create a situation where some of these dangers have the property that it doesn’t matter if three out of five of the companies are being safe, if the other two are being unsafe, it creates this negative externality. And I think the lack of uniformity is not fair to those of us who have put a lot of effort into being very thoughtful about these procedures. The second thing is I don’t think you can trust these companies to adhere to these voluntary plans on their own. Right? I like to think that Anthropic will, we do everything we can that we will, our RSP is checked by our long-term benefit trust, so we do everything we can to adhere to our own RSP.

Dario Amodei1:22:07

但你常听到各种关于不同公司的说法,比如‘哦,他们声称会提供这么多算力,结果并未兑现’,或‘他们承诺要做这件事,却并未做到’。我认为,逐一追究各家公司具体行为并无意义;但我坚信一个更根本的原则:倘若无人监督这些公司,倘若整个行业都缺乏监督机制,那么我们就无法保证自己会做正确的事,而当前的风险之高,容不得半点闪失。因此,我认为确立一项全行业统一遵循的标准至关重要,也必须确保整个行业切实履行多数业内主体早已公开申明其重要性、并明确承诺必将落实的事项。

英文原文

But you hear lots of things about various companies saying, “Oh, they said they would give this much compute and they didn’t. They said they would do this thing and the didn’t.” I don’t think it makes sense to litigate particular things that companies have done, but I think this broad principle that if there’s nothing watching over them, if there’s nothing watching over us as an industry, there’s no guarantee that we’ll do the right thing and the stakes are very high. And so I think it’s important to have a uniform standard that everyone follows and to make sure that simply that the industry does what a majority of the industry has already said is important and has already said that they definitely will do.

Dario Amodei1:22:52

没错,有些人——我认为存在一类人,他们原则上反对监管。我理解这种立场的来源。如果你前往欧洲,看到像GDPR这样的法规,或他们推行的其他一些举措,其中部分确实有益,但另一些则明显过于繁重、不必要,甚至可以说确实拖慢了创新步伐。因此,人们基于先验立场持此观点,是完全可以理解的。我明白他们为何从这一立场出发。但再次强调,我认为人工智能有所不同。如果我们回看几分钟前刚刚谈到的那些极其严峻的自主性与滥用风险,我认为这些风险非同寻常,理应引发非同寻常的强力应对。因此,我认为这一点至关重要。

英文原文

Right, some people, I think there’s a class of people who are against regulation on principle. I understand where that comes from. If you go to Europe and you see something like GDPR, you see some of the other stuff that they’ve done. Some of it’s good, but some of it is really unnecessarily burdensome and I think it’s fair to say really has slowed innovation. And so I understand where people are coming from on priors. I understand why people start from that position. But again, I think AI is different. If we go to the very serious risks of autonomy and misuse that I talked about just a few minutes ago, I think that those are unusual and they warrant an unusually strong response. And so I think it’s very important.

Dario Amodei1:23:44

同样,我们需要一项能获得普遍支持的方案。我认为SB 1047法案——尤其是其最初版本——存在的一个问题在于:它虽纳入了RSP的诸多结构要素,但也夹杂了不少笨拙条款,或会带来大量负担、繁琐手续,甚至可能偏离靶心,未能真正切中风险要害。你在推特上几乎听不到关于该法案具体内容的讨论,只看到人们笼统地为‘任何监管’欢呼;而反对者则常常编造出一些往往极不诚实的论点,例如称该法案将导致企业迁出加州(事实上该法案不适用于总部设在加州的企业)、称该法案仅适用于在加州开展业务的企业、称其将损害开源生态,或称其将引发诸如此类的一系列后果。

英文原文

Again, we need something that everyone can get behind. I think one of the issues with SB 1047, especially the original version of it was it had a bunch of the structure of RSPs, but it also had a bunch of stuff that was either clunky or that just would’ve created a bunch of burdens, a bunch of hassle and might even have missed the target in terms of addressing the risks. You don’t really hear about it on Twitter, you just hear about kind of people are cheering for any regulation. And then the folks who are against make up these often quite intellectually dishonest arguments about how it’ll make us move away from California, bill doesn’t apply if you’re headquartered in California, bill only applies if you do business in California, or that it would damage the open source ecosystem or that it would cause all of these things.

Dario Amodei1:24:43

我认为上述论点大多纯属无稽之谈,但确实存在更有力的反监管理由。有一位名叫迪恩·鲍尔(Dean Ball)的人,我认为他是一位非常严谨的学者型分析人士,专门研究监管一旦落地,可能如何自行演化、或因设计不当而产生何种后果。因此,我们一贯的立场是:我们确实认为该领域需要监管,但我们希望成为推动监管走向精准化的力量——即监管须聚焦于真正严峻的风险,并且具备可操作性,让业界能够切实遵行。因为我认为,监管倡导者尚未充分意识到一点:倘若我们最终出台一项目标模糊、定位不准的监管措施,徒然耗费大量人力物力,结果只会是人们纷纷抱怨:‘瞧,这些安全风险纯属无稽之谈。我刚被迫雇了十名律师填表,又不得不为一项显然毫无危险性的事项跑遍所有测试流程。’

英文原文

I think those were mostly nonsense, but there are better arguments against regulation. There’s one guy, Dean Ball, who’s really, I think, a very scholarly analyst who looks at what happens when a regulation is put in place in ways that they can kind of get a life of their own or how they can be poorly designed. And so our interest has always been we do think there should be regulation in this space, but we want to be an actor who makes sure that that regulation is something that’s surgical, that’s targeted at the serious risks and is something people can actually comply with. Because something I think the advocates of regulation don’t understand as well as they could is if we get something in place that’s poorly targeted, that wastes a bunch of people’s time, what’s going to happen is people are going to say, “See, these safety risks, this is nonsense. I just had to hire 10 lawyers to fill out all these forms. I had to run all these tests for something that was clearly not dangerous.”

Dario Amodei1:25:51

如此持续六个月之后,必将掀起一股声势浩大的民意浪潮,最终形成一种持久稳固的反监管共识。因此,我认为,真正渴望切实问责的人士所面临的最大敌人,恰恰是设计拙劣的监管。我们必须真正把事情做对。如果我能向监管倡导者说一句话,那就是:我希望他们更深入地理解这一动态机制;我们必须格外审慎,必须与那些真正拥有监管实践观察经验的人士展开对话。而那些亲历过监管落地过程的人,深知务必慎之又慎。倘若这只是一个次要问题,我或许会彻底反对监管。

英文原文

And after six months of that, there will be a ground swell and we’ll end up with a durable consensus against regulation. And so I think the worst enemy of those who want real accountability is badly designed regulation. We need to actually get it right. And if there’s one thing I could say to the advocates, it would be that I want them to understand this dynamic better and we need to be really careful and we need to talk to people who actually have experience seeing how regulations play out in practice. And the people who have seen that, understand to be very careful. If this was some lesser issue, I might be against regulation at all.

Dario Amodei1:26:32

但我想让反对者理解的是:其背后的根本性问题确实十分严峻。这些问题并非我本人或其他公司出于监管俘获动机而凭空捏造,也绝非科幻幻想,更不是其他任何一类虚妄之说。每次我们发布新模型,每隔几个月便会对这些模型的行为进行测评,结果发现它们在这些令人担忧的任务上的表现正持续提升,恰如其在有益、有价值、具经济实用性的任务上表现日益精进一样。因此,我真心希望——我认为SB 1047法案曾极具两极分化效应——一些最讲道理的反对者与最讲道理的支持者能够坐下来共同商议。在不同AI公司中,Anthropic是唯一一家以极为详尽的方式表达正面态度的AI公司;埃隆·马斯克(Elon Musk)曾在推特上简短地发表过一条积极评论,但谷歌、OpenAI、Meta、微软等几家巨头则基本持坚定反对立场。

英文原文

But what I want the opponents to understand is that the underlying issues are actually serious. They’re not something that I or the other companies are just making up because of regulatory capture, they’re not sci-fi fantasies, they’re not any of these things. Every time we have a new model, every few months we measure the behavior of these models and they’re getting better and better at these concerning tasks just as they are getting better and better at good, valuable, economically useful tasks. And so I would just love it if some of the former, I think SB 1047 was very polarizing, I would love it if some of the most reasonable opponents and some of the most reasonable proponents would sit down together. And I think that the different AI companies, Anthropic was the only AI company that felt positively in a very detailed way. I think Elon tweeted briefly something positive, but some of the big ones like Google, OpenAI, Meta, Microsoft were pretty staunchly against.

Dario Amodei1:27:49

因此,我真诚期盼:关键利益相关方中,最具思辨能力的支持者与最具思辨能力的反对者,能够坐在一起探讨:我们该如何设计一套解决方案,既能让支持者切实感受到风险显著降低,又能让反对者确信其不会对产业或创新造成超出必要限度的阻碍?不知为何,此事已过度两极化,致使这两类群体未能如本应那样展开务实对话。我深感紧迫。我真心认为我们必须在2025年内有所行动。倘若到2025年底我们依然无所作为,我将深感忧虑。目前我尚不忧虑,因为正如前述,这些风险尚未真正显现;但我认为,时间已然所剩无几。

英文原文

So I would really is if some of the key stakeholders, some of the most thoughtful proponents and some of the most thoughtful opponents would sit down and say how do we solve this problem in a way that the proponents feel brings a real reduction in risk and that the opponents feel that it is not hampering the industry or hampering innovation any more necessary than it needs to. I think for whatever reason, that things got too polarized and those two groups didn’t get to sit down in the way that they should. And I feel urgency. I really think we need to do something in 2025. If we get to the end of 2025 and we’ve still done nothing about this, then I’m going to be worried. I’m not worried yet because, again, the risks aren’t here yet, but I think time is running short.

Lex Fridman1:28:44

并提出一项精准施策的方案,正如您刚才所言。

英文原文

And come up with something surgical, like you said.

Dario Amodei1:28:46

是的,是的,正是如此。我们必须摆脱当前这种‘极端推崇安全’与‘极端反对监管’的激烈对立话语。这种对立已演变为推特上的骂战,而此类争斗绝不会带来任何积极成果。

英文原文

Yeah, yeah, yeah, exactly. And we need to get away from this intense pro safety versus intense anti-regulatory rhetoric. It’s turned into these flame wars on Twitter and nothing good’s going to come of that.

Lex Fridman1:29:04

外界对这一领域内各方参与者充满好奇。其中一位元老级玩家便是OpenAI。您曾在OpenAI积累了多年经验。能否请您谈谈您在那里经历的故事与历史?

英文原文

So there’s a lot of curiosity about the different players in the game. One of the OGs is OpenAI. You’ve had several years of experience at OpenAI. What’s your story and history there?

Dario Amodei1:29:14

是的。我在OpenAI工作了大约五年。最后两年左右,我担任该公司的研究副总裁。大概是我和伊利亚·苏茨克弗(Ilya Sutskever)两人共同确立了公司的研究方向。大约在2016年或2017年,我首次真正开始相信、或者说进一步确认了我对‘缩放假说’(scaling hypothesis)的信念——当时伊利亚曾对我留下一句著名的话:‘你必须明白,这些模型的本质就是渴望学习。模型本身,就是渴望学习。’有时,你会听到某句短短一句话、某个顿悟时刻,瞬间令你豁然开朗:‘啊,这就解释了一切!它解释了我此前目睹过的上千种现象。’自此以后,我脑海中便始终浮现出这样一幅图景:只要你以正确方式优化模型、以正确方向引导模型,它们就天然渴望学习,天然渴望解决问题——无论问题本身是什么。

英文原文

Yeah. So I was at OpenAI for roughly five years. For the last, I think it was couple years, I was vice president of research there. Probably myself and Ilya Sutskever were the ones who really kind of set the research direction. Around 2016 or 2017, I first started to really believe in or at least confirm my belief in the scaling hypothesis when Ilya famously said to me, “The thing you need to understand about these models is they just want to learn. The models just want to learn.” And again, sometimes there are these one sentences, these then cones, that you hear them and you’re like, “Ah, that explains everything. That explains a thousand things that I’ve seen.” And then ever after, I had this visualization in my head of you optimize the models in the right way, you point the models in the right way, they just want to learn. They just want to solve the problem regardless of what the problem is.

Lex Fridman1:30:08

所以,本质上就是别挡它们的道?

英文原文

So get out of their way, basically?

Dario Amodei1:30:10

别挡它们的道。没错。

英文原文

Get out of their way. Yeah.

Lex Fridman1:30:11

好的。

英文原文

Okay.

Dario Amodei1:30:11

不要强加你自己关于他们应该如何学习的想法。这与理查德·萨顿(Rich Sutton)提出的‘苦涩的教训’(the bitter lesson),或格温(Gwern)提出的‘扩展假说’(scaling hypothesis)所表达的观点如出一辙。总体而言,我的思路是受到伊利亚(Ilya)以及亚历克·拉德福德(Alec Radford)等人的启发——他主导了最初的GPT-1,并在此基础上全力推进;而我与我的合作者则聚焦于GPT-2、GPT-3,以及‘基于人类反馈的强化学习’(RL from Human Feedback),后者旨在应对早期的安全性与稳健性问题,例如辩论(debate)、放大(amplification)等方法,且高度重视可解释性(interpretability)。因此,再次强调,这是安全性与规模化(scaling)的结合。大概在2018年、2019年和2020年这几年间,我和我的合作者——其中许多人后来成为Anthropic的联合创始人——形成了清晰的愿景,并主导了这一方向的发展。

英文原文

Don’t impose your own ideas about how they should learn. And this was the same thing as Rich Sutton put out in the bitter lesson or Gwern put out in the scaling hypothesis. I think generally the dynamic was I got this kind of inspiration from Ilya and from others, folks like Alec Radford, who did the original GPT-1 and then ran really hard with it, me and my collaborators, on GPT-2, GPT-3, RL from Human Feedback, which was an attempt to kind of deal with the early safety and durability, things like debate and amplification, heavy on interpretability. So again, the combination of safety plus scaling. Probably 2018, 2019, 2020, those were kind of the years when myself and my collaborators, probably many of whom became co-founders of Anthropic, kind of really had a vision and drove the direction.

Lex Fridman1:31:11

你为什么离开?你为何决定离开?

英文原文

Why’d you leave? Why’d you decide to leave?

Dario Amodei1:31:13

是的,我这么来说吧——我认为这与‘向上竞争’(race to the top)密切相关。在我任职OpenAI期间,我逐渐认识到并愈发重视‘扩展假说’,同时也愈发意识到:在推进扩展假说的同时,安全性同样至关重要。第一点,即扩展假说,OpenAI当时已开始接纳;第二点,即安全性,某种意义上始终是OpenAI对外传达信息的一部分。但在我在那里工作的多年间,我对如何处理这些问题、如何将这些技术推向世界、组织应秉持何种原则,形成了自己独特的构想。当然,我们曾就‘公司是否该这么做’‘公司是否该那么做’展开过大量讨论。外界存在大量不实信息。

英文原文

Yeah, so look, I’m going to put things this way and I think it ties to the race to the top, which is in my time at OpenAI, what I come to see as I’d come to appreciate the scaling hypothesis and as I’d come to appreciate kind of the importance of safety along with the scaling hypothesis. The first one I think OpenAI was getting on board with. The second one in a way had always been part of OpenAI’s messaging. But over many years of the time that I spent there, I think I had a particular vision of how we should handle these things, how we should be brought out in the world, the kind of principles that the organization should have. And look, there were many, many discussions about should the company do this, should the company do that? There’s a bunch of misinformation out there.

Dario Amodei1:32:07

有人说我们离开是因为不满与微软的交易——这是错误的。尽管当时确实围绕与微软合作的具体方式展开了大量讨论,提出了大量疑问。也有人说我们离开是因为反对商业化——这也不对。我们开发了GPT-3,而它正是被商业化的模型;我也参与了商业化工作。真正的问题在于:究竟该如何商业化?人类文明正沿着这条路径走向极其强大的人工智能。那么,什么才是审慎、坦率、诚实的方式?怎样才能建立公众对组织及个体的信任?我们如何从当下抵达那个理想状态?我们如何真正形成一套切实可行的、确保做对事情的远见?安全性不应仅仅是一句用于招聘宣传的空话。我认为,归根结底,如果你对此拥有自己的远见,那就无需在意他人的远见。

英文原文

People say we left because we didn’t like the deal with Microsoft. False. Although, it was like a lot of discussion, a lot of questions about exactly how we do the deal with Microsoft. We left because we didn’t like commercialization. That’s not true. We built GPD-3, which was the model that was commercialized. I was involved in commercialization. It’s more, again, about how do you do it? Civilization is going down this path to very powerful AI. What’s the way to do it? That is cautious, straightforward, honest, that builds trust in the organization and in individuals. How do we get from here to there and how do we have a real vision for how to get it right? How can safety not just be something we say because it helps with recruiting. And I think at the end of the day, if you have a vision for that, forget about anyone else’s vision.

Dario Amodei1:33:01

我不想谈论任何他人的远见。如果你对如何实现目标已有清晰远见,你就应该付诸行动,去践行这一远见。试图与他人争辩其远见,是极其低效的。你或许认为对方做法不对,或许怀疑对方不够坦诚——谁知道呢?也许你是对的,也许你错了。但你应该做的,是召集一批你信任的人,共同出发,将你的远见变为现实。如果你的远见足够有说服力,如果你能让它打动人心——无论是从伦理层面还是市场层面,如果你能创建一家人们愿意加入的公司,践行被公众视为合理的行为准则,同时还能在生态系统中维持自身地位——那么,其他人就会效仿你。

英文原文

I don’t want to talk about anyone else’s vision. If you have a vision for how to do it, you should go off and you should do that vision. It is incredibly unproductive to try and argue with someone else’s vision. You might think they’re not doing it the right way. You might think they’re dishonest. Who knows? Maybe you’re right, maybe you’re not. But what you should do is you should take some people you trust and you should go off together and you should make your vision happen. And if your vision is compelling, if you can make it appeal to people, some combination of ethically in the market, if you can make a company that’s a place people want to join, that engages in practices that people think are reasonable while managing to maintain its position in the ecosystem at the same time, if you do that, people will copy it.

Dario Amodei1:33:52

而你正在践行这一理念的事实——尤其是你比别人做得更好这一事实——会以一种远比你作为下属与老板争辩更有力的方式,促使他们改变自身行为。我无法再比这说得更具体了,但我认为,试图让别人的远见看起来像你的远见,总体上是极低效的。更高效的做法是独立开展一项干净利落的实验,明确宣告:‘这是我们的远见,这是我们做事的方式。你们的选择只有三种:无视我们、拒绝我们所做之事,或者开始变得越来越像我们。’模仿是最真诚的恭维。这种效应会体现在客户行为上,体现在公众反应上,也体现在人们择业倾向上。再次强调,最终的关键并非某一家公司胜出、另一家公司落败。

英文原文

And the fact that you are doing it, especially the fact that you’re doing it better than they are, causes them to change their behavior in a much more compelling way than if they’re your boss and you’re arguing with them. I don’t know how to be any more specific about it than that, but I think it’s generally very unproductive to try and get someone else’s vision to look like your vision. It’s much more productive to go off and do a clean experiment and say, “This is our vision, this is how we’re going to do things. Your choice is you can ignore us, you can reject what we’re doing or you can start to become more like us.” And imitation is the sincerest form of flattery. And that plays out in the behavior of customers, that plays out in the behavior of the public, that plays out in the behavior of where people choose to work. And again, at the end, it’s not about one company winning or another company winning.

Dario Amodei1:34:48

如果我们或另一家公司推行某种实践,而人们真心觉得这种实践颇具吸引力——我强调的是实质内容,而非表面功夫;我认为研究人员相当成熟老练,他们关注的是实质——随后其他公司开始效仿这种实践,因而获得成功。这很好,这就是成功,这就是‘向上竞争’。最终谁胜出并不重要,只要所有公司都在相互效仿彼此的良好实践即可。我常这样理解:我们所有人真正恐惧的是‘向下竞争’(race to the bottom),而‘向下竞争’中谁胜出都无关紧要,因为所有人都将失败。在最极端的情形下,我们造出完全自主的人工智能,机器人奴役人类之类——这半开玩笑,但确属可能发生的最极端后果。届时,哪家公司领先已毫无意义。反之,若我们促成一场‘向上竞争’,让各公司竞相践行良好实践,那么最终谁胜出、甚至谁率先发起这场‘向上竞争’,都不再重要。

英文原文

If we or another company are engaging in some practice that people find genuinely appealing, and I want it to be in substance, not just an appearance and I think researchers are sophisticated and they look at substance, and then other companies start copying that practice and they win because they copied that practice. That’s great. That’s success. That’s like the race to the top. It doesn’t matter who wins in the end as long as everyone is copying everyone else’s good practices. One way I think of it is the thing we’re all afraid of is the race to the bottom and the race to the bottom doesn’t matter who wins because we all lose. In the most extreme world, we make this autonomous AI that the robots enslave us or whatever. That’s half joking, but that is the most extreme thing that could happen. Then it doesn’t matter which company was ahead. If instead you create a race to the top where people are competing to engage in good practices, then at the end of the day, it doesn’t matter who ends up winning, it doesn’t even matter who started the race to the top.

Dario Amodei1:35:57

关键不在于标榜美德,而在于推动整个系统进入一个比此前更优的均衡状态。单个公司可在其中发挥一定作用:可以助其启动,也可以助其加速。坦率地说,我认为其他公司的个人也已如此行动。比如,当我们发布《负责任的斯卡利普特》(RSP)时,其他公司的从业者会受此激励,更加努力地推动自家公司落实类似举措;有时其他公司采取的行动甚至让我们感叹:‘哦,这是项好实践,我们认为不错,我们也该采纳。’唯一的区别在于,我们力求更主动进取:当他人首创某项实践时,我们力求更快、更广泛地采纳。但我认为,这种动态机制才是我们真正应关注的焦点;它能抽离出‘哪家公司胜出’‘谁信任谁’等具体问题。我认为所有这些戏剧性纷争都极度乏味,真正重要的是我们所有人所处的生态系统,以及如何让这个生态系统变得更好——因为正是这个系统约束着所有参与者。

英文原文

The point isn’t to be virtuous, the point is to get the system into a better equilibrium than it was before. And individual companies can play some role in doing this. Individual companies can help to start it, can help to accelerate it. And frankly, I think individuals at other companies have done this as well. The individuals that when we put out an RSP react by pushing harder to get something similar done at other companies, sometimes other companies do something that’s we’re like, “Oh, it’s a good practice. We think that’s good. We should adopt it too.” The only difference is I think we try to be more forward leaning. We try and adopt more of these practices first and adopt them more quickly when others invent them. But I think this dynamic is what we should be pointing at and that I think it abstracts away the question of which company’s winning, who trusts who. I think all these questions of drama are profoundly uninteresting and the thing that matters is the ecosystem that we all operate in and how to make that ecosystem better because that constrains all the players.

Lex Fridman1:37:06

那么,Anthropic就是这样一个基于‘人工智能安全性究竟应具象为何种形态’这一具体理念而构建的干净利落的实验?

英文原文

And so Anthropic is this kind of clean experiment built on a foundation of what concretely AI safety should look like?

Dario Amodei1:37:13

嗯,我确信我们在过程中犯过不少错误。完美的组织根本不存在。它必须应对一千名员工各自的不完美;必须应对包括我在内的领导层的不完美;还必须应对董事会、长期受益信托(long-term benefit trust)等负责监督领导层之不完美的人员的不完美。这本质上是一群不完美的人,以不完美的方式,努力趋近某个永远无法被完美实现的理想目标。这正是你选择加入时所承诺承担的,也将永远如此。

英文原文

Well, look, I’m sure we’ve made plenty of mistakes along the way. The perfect organization doesn’t exist. It has to deal with the imperfection of a thousand employees. It has to deal with the imperfection of our leaders, including me. It has to deal with the imperfection of the people we’ve put to oversee the imperfection of the leaders like the board and the long-term benefit trust. It’s all a set of imperfect people trying to aim imperfectly at some ideal that will never perfectly be achieved. That’s what you sign up for. That’s what it will always be.

Dario Amodei1:37:45

但‘不完美’并不意味着就此放弃。总存在‘更好’与‘更差’之分。我们希望自身表现足够出色,从而开始建立一些全行业都能采纳的实践。我推测,多家此类公司将取得成功:Anthropic会成功,我过去供职过的那些公司也会成功,其中一些公司或许会比另一些更成功。但相较之下,更重要的是我们能否协调整个行业的激励机制。而这部分通过‘向上竞争’实现,部分通过RSP等机制实现,部分则通过有针对性的精准监管实现。

英文原文

But imperfect doesn’t mean you just give up. There’s better and there’s worse. And hopefully, we can do well enough that we can begin to build some practices that the whole industry engages in. And then my guess is that multiple of these companies will be successful. Anthropic will be successful. These other companies, like ones I’ve been at the past, will also be successful. And some will be more successful than others. That’s less important than, again, that we align the incentives of the industry. And that happens partly through the race to the top, partly through things like RSP, partly through, again, selected surgical regulation.

Lex Fridman1:38:25

你提到‘人才密度胜过人才规模’,能否解释一下?能否进一步展开说明?

英文原文

You said talent density beats talent mass, so can you explain that? Can you expand on that?

Dario Amodei1:38:25

是的。

英文原文

Yeah.

Lex Fridman1:38:31

能否谈谈,打造一支顶尖的人工智能研究员与工程师团队,究竟需要什么?

英文原文

Can you just talk about what it takes to build a great team of AI researchers and engineers?

Dario Amodei1:38:37

这是那种‘每月都显得更真实’的论断之一。每个月我都会觉得这句话比上个月更真实。所以,假如我们做一个思想实验:假设你有一支由100人组成的团队,个个极其聪明、充满干劲,且高度认同公司使命——这就是你的公司;或者你也可以拥有一支千人团队,其中200人极其聪明、高度认同使命,而另外800人,我们就姑且说,是从大型科技公司里随机挑选出来的普通员工——你更愿意选择哪一种?从人才总量来看,千人团队显然更多;你甚至拥有数量更多的、极其优秀、高度认同使命、极其聪明的人才。但问题在于:每当一位极其优秀的人才环顾四周,看到的都是同样优秀、同样专注投入的同事时,这种氛围就为整个组织定下了基调——它让每个人都深受鼓舞,渴望在同一个地方工作;也让每个人彼此信任。

英文原文

This is one of these statements that’s more true every month. Every month I see this statement as more true than I did the month before. So if I were to do a thought experiment, let’s say you have a team of 100 people that are super smart, motivated and aligned with the mission and that’s your company. Or you can have a team of a thousand people where 200 people are super smart, super aligned with the mission and then 800 people are, let’s just say you pick 800 random big tech employees, which would you rather have? The talent mass is greater in the group of a thousand people. You have even a larger number of incredibly talented, incredibly aligned, incredibly smart people. But the issue is just that if every time someone super talented looks around, they see someone else super talented and super dedicated, that sets the tone for everything. That sets the tone for everyone is super inspired to work at the same place. Everyone trusts everyone else.

Dario Amodei1:39:42

如果你的团队规模达到一千人甚至一万人,而组织状态已明显退化,你就无法再进行有效筛选,只能随机招人。这时,你就不得不设置大量流程和诸多管控机制,原因仅仅在于人们彼此之间缺乏充分信任,或必须不断裁决各种政治性纷争。诸如此类的因素,严重拖慢了组织的运转效率。目前我们团队人数已接近一千人,我们一直努力确保这千人中尽可能高的比例,都是极其优秀、技能超群的人才。这也是过去几个月我们大幅放缓招聘节奏的原因之一。今年前七、八个月,我们团队规模大概从300人增长到了800人;而如今增速已明显放缓——最近三个月,我们仅从800人增至约900人、950人左右。具体数字请勿直接引用,但我认为,千人规模附近存在一个拐点,我们希望在此之后更加审慎地规划扩张路径。

英文原文

If you have a thousand or 10,000 people and things have really regressed, you are not able to do selection and you’re choosing random people, what happens is then you need to put a lot of processes and a lot of guardrails in place just because people don’t fully trust each other or you have to adjudicate political battles. There are so many things that slow down the org’s ability to operate. And so we’re nearly a thousand people and we’ve tried to make it so that as large a fraction of those thousand people as possible are super talented, super skilled, it’s one of the reasons we’ve slowed down hiring a lot in the last few months. We grew from 300 to 800, I believe, I think in the first seven, eight months of the year and now we’ve slowed down. The last three months, we went from 800 to 900, 950, something like that. Don’t quote me on the exact numbers, but I think there’s an inflection point around a thousand and we want to be much more careful how we grow.

Dario Amodei1:40:42

早期乃至现在,我们都招募了大量物理学家。理论物理学家学习新知识的速度极快。而近来,在持续沿用这一招聘策略的同时,我们在研究岗与软件工程岗均设定了极高门槛,招入了许多资深人才,包括此前曾在该领域其他公司任职的人员;我们始终保持着极为严苛的筛选标准。若不加留意,企业很容易从百人规模迅速膨胀至千人,再跃升至万人,却未能确保全员目标一致。这种统一目标的力量极为强大。倘若一家公司内部充斥着众多各自为政的‘诸侯领地’,每一块都想按自己的逻辑行事、各自优化自身指标,那么几乎什么事都难以推进。但若每位员工都能看清公司的整体使命,彼此间充满信任,并坚定致力于做正确的事,那便是一种‘超能力’——仅凭这一点,我认为就足以克服几乎所有其他劣势。

英文原文

Early on and now as well, we’ve hired a lot of physicists. Theoretical physicists can learn things really fast. Even more recently, as we’ve continued to hire that, we’ve really had a high bar on both the research side and the software engineering side, have hired a lot of senior people, including folks who used to be at other companies in this space, and we’ve just continued to be very selective. It’s very easy to go from a hundred to a thousand, a thousand to 10,000 without paying attention to making sure everyone has a unified purpose. It’s so powerful. If your company consists of a lot of different fiefdoms that all want to do their own thing, they’re all optimizing for their own thing, it’s very hard to get anything done. But if everyone sees the broader purpose of the company, if there’s trust and there’s dedication to doing the right thing, that is a superpower. That in itself I think can overcome almost every other disadvantage.

Lex Fridman1:41:41

对史蒂夫·乔布斯而言,就是‘A级人才’。所谓A级人才,就是希望环顾四周时看到的也全是A级人才——换种说法而已。

英文原文

And to Steve Jobs, A players. A players want to look around and see other A players is another way of saying that.

Dario Amodei1:41:42

没错。

英文原文

Correct.

Lex Fridman1:41:48

我不知道这背后反映的是何种人性本质,但目睹那些并未全身心执着于单一使命的人,确实令人泄气;而反过来看,见到这样执着的人,则极具激励作用。这很有意思。那么,根据你与众多杰出人士共事的经验,要成为一名出色的AI研究员或工程师,究竟需要什么?

英文原文

I don’t know what that is about human nature, but it is demotivating to see people who are not obsessively driving towards a singular mission. And it is on the flip side of that, super motivating to see that. It’s interesting. What’s it take to be a great AI researcher or engineer from everything you’ve seen from working with so many amazing people?

Dario Amodei1:42:09

是的。我认为首要品质——尤其在研究领域,但其实对工程领域也同样适用——是开放心态。听起来‘开放心态’似乎很容易做到,对吧?你只需说一句‘哦,我对任何事物都持开放态度’就行。但回顾我自己早年关于‘缩放假说’的经历:我所看到的数据,和其他人并无二致。我不认为自己在编程能力或提出研究构想方面,比曾共事过的数百人更出色;某种程度上,我甚至更逊色——比如精准定位程序缺陷、编写GPU内核代码这类工作,我从未擅长过。我可以随口指出这里上百位同事,他们在这些方面都远胜于我。

英文原文

Yeah. I think the number one quality, especially on the research side, but really both, is open mindedness. Sounds easy to be open-minded, right? You’re just like, “Oh, I’m open to anything.” But if I think about my own early history in this scaling hypothesis, I was seeing the same data others were seeing. I don’t think I was a better programmer or better at coming up with research ideas than any of the hundreds of people that I worked with. In some ways, I was worse. I’ve never precise programming of finding the bug, writing the GPU kernels. I could point you to a hundred people here who are better at that than I am.

Dario Amodei1:42:53

但我想,我真正与众不同的地方在于:我愿意以全新的视角审视事物。当别人说‘哦,我们还没找到合适的算法’‘我们尚未摸索出正确的实现方式’时,我只会想:‘哦,不确定呢。这个神经网络有3000万个参数,如果把它增加到5000万个会怎样?咱们画几张图看看吧。’这种最基础的科学思维模式就是:‘哇,我发现了一个可调节的变量,改变它会发生什么?咱们试试不同取值,画张图出来。’哪怕是最简单不过的操作——比如调整参数数量——这根本不是博士级别的实验设计,纯粹简单又朴素。只要有人告诉你这件事很重要,任何人都能照做。理解起来也不难,你根本不需要多么聪明就能想到这一点。

英文原文

But the thing that I think I did have that was different was that I was just willing to look at something with new eyes. People said, “Oh, we don’t have the right algorithms yet. We haven’t come up with the right way to do things.” And I was just like, “Oh, I don’t know. This neural net has 30 million parameters. What if we gave it 50 million instead? Let’s plot some graphs.” That basic scientific mindset of like, “Oh man,” I see some variable that I could change. What happens when it changes? Let’s try these different things and create a graph. For even, this was the simplest thing in the world, change the number of, this wasn’t PhD level experimental design, this was simple and stupid. Anyone could have done this if you just told them that it was important. It’s also not hard to understand. You didn’t need to be brilliant to come up with this.

Dario Amodei1:43:54

但当你把这两点结合起来,就会发现仅有极少数人——个位数级别的人——通过意识到这一点,推动了整个领域的发展。历史上的重大发现往往也是如此。因此,这种开放心态,以及这种愿以全新视角观察事物的意愿(而这往往源于初入该领域的新鲜感;事实上,经验有时反而会成为这种特质的障碍),才是最重要的素质。它极难识别与评估,但我认为它至关重要,因为一旦你发现了某种真正新颖的思维方式,并主动付诸实践,其带来的变革将是彻底而深远的。

英文原文

But you put the two things together and some tiny number of people, some single digit number of people have driven forward the whole field by realizing this. And it’s often like that. If you look back at the discoveries in history, they’re often like that. And so this open-mindedness and this willingness to see with new eyes that often comes from being newer to the field, often experience is a disadvantage for this, that is the most important thing. It’s very hard to look for and test for, but I think it’s the most important thing because when you find something, some really new way of thinking about things, when you have the initiative to do that, it’s absolutely transformative.

Lex Fridman1:44:34

此外,还需具备快速开展实验的能力;面对实验结果,保持开放、好奇的心态,以崭新的眼光审视数据,去探究数据本身究竟在传达什么信息——这在机制可解释性(mechanistic interpretability)研究中尤为适用。

英文原文

And also be able to do kind of rapid experimentation and, in the face of that, be open-minded and curious and looking at the data with these fresh eyes and seeing what is it that it’s actually saying. That applies in mechanistic interpretability.

Dario Amodei1:44:46

这又是另一个例证。机制可解释性领域的早期工作极其简单,只是此前根本没人想过要关注这个问题罢了。

英文原文

It’s another example of this. Some of the early work and mechanistic interpretability so simple, it’s just no one thought to care about this question before.

Lex Fridman1:44:56

你刚才谈到了成为优秀AI研究员所需具备的素质。我们能否把时间拨回一点,谈谈你对有意投身AI领域的人有什么建议?他们很年轻,正满怀期待地思考:我该如何影响这个世界?

英文原文

You said what it takes to be a great AI researcher. Can we rewind the clock back, what advice would you give to people interested in AI? They’re young, looking…

Lex Fridman1:45:00

你对有意投身AI领域的人有什么建议?他们很年轻,正满怀期待地思考:我该如何影响这个世界?

英文原文

What advice would you give to people interested in AI? They’re young. Looking forward to how can I make an impact on the world?

Dario Amodei1:45:06

我认为,我给出的首要建议就是:立刻动手玩转这些模型。实际上,我有点担心——如今这听起来似乎已是显而易见的建议了。但三年前并非如此,那时人们往往从‘哦,让我先读读最新的强化学习论文’开始。当然,你也应该这么做;但如今随着模型与API的普及程度大幅提升,人们已越来越多地采取这种实践方式。我认为,关键在于获得亲身体验。这些模型是全新的技术产物,尚无人真正完全理解,因此亲身操作、反复试错,积累实践经验至关重要。我还想再次强调:要勇于探索新方向、尝试新思路。当前仍存在大量未被开垦的领域。例如,机制可解释性仍处于非常早期的阶段;相比开发新型模型架构,投身于此可能更具价值——因为它虽已比从前更受关注,但从事该方向的研究者可能仅约百人,远未达到万人规模。

英文原文

I think my number one piece of advice is to just start playing with the models. Actually, I worry a little, this seems like obvious advice now. I think three years ago it wasn’t obvious and people started by, “Oh, let me read the latest reinforcement learning paper.” And you should do that as well, but now with wider availability of models and APIs, people are doing this more. But, I think just experiential knowledge. These models are new artifacts that no one really understands and so getting experience playing with them. I would also say again, in line with the do something new, think in some new direction, there are all these things that haven’t been explored. For example, mechanistic interpretability is still very new. It’s probably better to work on that than it is to work on new model architectures, because it’s more popular than it was before. There are probably 100 people working on it, but there aren’t like 10,000 people working on it.

Dario Amodei1:46:07

这是一个极具潜力的研究沃土,遍地都是‘低垂的果实’,你只需信步走过,随手便可摘取。不知为何,人们对它的兴趣尚不够浓厚。我认为,长周期学习与长周期任务相关领域也大有可为。在评测方法方面,我们的研究仍处于非常初级的阶段,尤其是针对在现实世界中动态运行的系统。多智能体(multi-agent)方向也存在不少值得探索的空间。我的建议是:‘滑向冰球将要到达的位置’(Skate where the puck is going)。你无需多么聪明才能想到这一点。所有五年后将令人兴奋的方向,如今甚至已被奉为常识;但不知为何,总存在一道无形屏障,阻碍人们全力投入,或让人畏惧去做那些尚未成为主流的事情。我不知道这种现象为何发生,但突破这道屏障,正是我给出的首要建议。

英文原文

And it’s just this fertile area for study. There’s so much low-hanging fruit, you can just walk by and you can pick things. For whatever reason, people aren’t interested in it enough. I think there are some things around long horizon learning and long horizon tasks, where there’s a lot to be done. I think evaluations, we’re still very early in our ability to study evaluations, particularly for dynamic systems acting in the world. I think there’s some stuff around multi-agent. Skate where the puck is going is my advice, and you don’t have to be brilliant to think of it. All the things that are going to be exciting in five years, people even mention them as conventional wisdom, but it’s just somehow there’s this barrier that people don’t double down as much as they could, or they’re afraid to do something that’s not the popular thing. I don’t know why it happens, but getting over that barrier, that’s my number one piece of advice.

Lex Fridman1:47:14

让我们稍微聊聊后训练(post-training)吧。目前主流的后训练流程似乎包罗万象:监督微调(supervised fine-tuning)、基于人类反馈的强化学习(RLHF)、宪法式人工智能(constitutional AI)及其衍生的基于AI反馈的强化学习(RLAIF)……

英文原文

Let’s talk if we could a bit about post-training. So it seems that the modern post-training recipe has a little bit of everything. So supervised fine-tuning, RLHF, the constitutional AI with RLAIF-

Dario Amodei1:47:32

最佳缩写词。

英文原文

Best acronym.

Lex Fridman1:47:33

又是那个命名问题。然后是合成数据——看起来大量使用了合成数据,或者至少是在努力探索如何生成高质量的合成数据。所以,如果这确实是让Anthropic模型如此出色的核心秘诀,那么其中多少‘魔力’来自预训练?又有多少来自后训练?

英文原文

It’s the, again, that naming thing. And then synthetic data. Seems like a lot of synthetic data, or at least trying to figure out ways to have high quality synthetic data. So if this is a secret sauce that makes Anthropic clause so incredible, how much of the magic is in the pre-training? How much of it is in the post-training?

Dario Amodei1:47:54

是的。首先,我们自身其实并不完全能准确衡量这一点。当你看到模型展现出某种出色的特性时,有时很难判断它究竟源自预训练还是后训练。我们开发了一些方法来尝试区分这两者,但这些方法并不完美。其次,我想说的是:当存在某种优势时——我认为我们在强化学习(RL)方面总体上一直做得相当好,或许甚至可以说是最好的——尽管我并不确定,毕竟我看不到其他公司内部的具体情况。通常这种优势并非‘天啊,我们掌握了一种别人没有的神秘魔法方法’,而更常是‘嗯,我们的基础设施更完善,因此能运行更长时间’,或‘我们获得了更高品质的数据’,或‘我们能更好地过滤数据’,又或‘我们成功融合了多种方法并持续实践’。

英文原文

Yeah. So first of all, we’re not perfectly able to measure that ourselves. When you see some great character ability, sometimes it’s hard to tell whether it came from pre-training or post-training. We developed ways to try and distinguish between those two, but they’re not perfect. The second thing I would say is, when there is an advantage and I think we’ve been pretty good in general at RL, perhaps the best, although I don’t know, I don’t see what goes on inside other companies. Usually it isn’t, “Oh my God, we have this secret magic method that others don’t have.” Usually it’s like, “Well, we got better at the infrastructure so we could run it for longer,” or, “We were able to get higher quality data,” or, “We were able to filter our data better, or “We were able to combine these methods and practice.”

Dario Amodei1:48:41

通常,这不过是一些枯燥的实践积累与行业技艺。因此,当我思考如何在模型训练方面做出特别之处时——不仅限于训练本身,更关键的是——我其实更多地将其类比为设计飞机或汽车:这并不仅仅是‘哦,天哪,我手握蓝图’那么简单;也许蓝图确实能帮你造出下一代飞机,但真正更重要的,是我们围绕设计过程所形成的那种文化性、技艺性的思维方式,它远比我们所能发明的任何具体‘小装置’都更为重要。

英文原文

It’s usually some boring matter of practice and trade craft. So when I think about how to do something special in terms of how we train these models both, but even more so I really think of it a little more, again, as designing airplanes or cars. It’s not just like, “Oh, man. I have the blueprint.” Maybe that makes you make the next airplane. But there’s some cultural trade craft of how we think about the design process that I think is more important than any particular gizmo we’re able to invent.

Lex Fridman1:49:17

好的。那我来问一个关于具体技术的问题:先从RLHF(基于人类反馈的强化学习)说起——仅凭直觉,甚至近乎哲学层面地看,你认为RLHF为何如此有效?

英文原文

Okay. Well, let me ask you about specific techniques. So first on RLHF, what do you think, just zooming out intuition, almost philosophy … Why do you think RLHF works so well?

Dario Amodei1:49:28

如果回溯到缩放假说(scaling hypothesis),其中一种应对方式是:如果你用X量级的资源进行训练,并投入足够算力,最终就能得到X级别的效果。而RLHF的优势恰恰在于,它能较好地使模型执行人类希望它做的事——更精确地说,是执行那些在短时间内观察模型、并对比不同可能回复后,人类所偏好的回复。但这套机制无论从安全性还是能力角度看都不完美:人类往往无法精准识别模型真正意图,且人类当下一时的偏好,未必等同于其长期真实所需。

英文原文

If I go back to the scaling hypothesis, one of the ways to skate the scaling hypothesis is, if you train for X and you throw enough compute at it, then you get X. And so RLHF is good at doing what humans want the model to do, or at least to state it more precisely doing what humans who look at the model for a brief period of time and consider different possible responses, what they prefer as the response, which is not perfect from both the safety and capabilities perspective, in that humans are often not able to perfectly identify what the model wants and what humans want in the moment may not be what they want in the long term.

Dario Amodei1:50:05

因此其中存在大量微妙之处;但模型确实擅长产出某种意义上‘人类浅层想要’的内容。事实上,你甚至无需为此投入过多算力,原因在于另一点:一个强大的预训练模型已‘半程抵达’任意目标。换言之,一旦拥有了预训练模型,你就已具备所有必要表征,足以将模型引导至你期望的位置。

英文原文

So there’s a lot of subtlety there, but the models are good at producing what the humans in some shallow sense want. And it actually turns out that you don’t even have to throw that much compute at it, because of another thing, which is this thing about a strong pre-trained model being halfway to anywhere. So once you have the pre-trained model, you have all the representations you need to get the model where you want it to go.

Lex Fridman1:50:32

那么,你认为RLHF是让模型变得更聪明了,还是仅仅让它在人类眼中显得更聪明?

英文原文

So do you think RLHF makes the model smarter, or just appear smarter to the humans?

Dario Amodei1:50:41

我不认为它让模型变得更聪明,也不认为它只是让模型在人类眼中显得更聪明。RLHF更像是在人类与模型之间架起一座桥梁。我完全可以拥有一个极其聪明却完全无法沟通的系统——我们都认识这类人:非常聪明,却让人听不懂他们在说什么。因此,我认为RLHF只是弥合了这一鸿沟。而且,它并非我们唯一采用的强化学习形式,也绝非未来唯一会采用的强化学习类型。我认为强化学习本身具备让模型变得更聪明、推理能力更强、运行表现更优、乃至发展出新技能的潜力;或许在某些情况下,这种提升甚至可通过人类反馈实现。但就目前我们所实施的RLHF而言,它大多尚未实现上述目标——尽管我们正迅速迈向这一阶段。

英文原文

I don’t think it makes the model smarter. I don’t think it just makes the model appear smarter. It’s like RLHF bridges the gap between the human and the model. I could have something really smart that can’t communicate at all. We all know people like this, people who are really smart but can’t understand what they’re saying. So I think RLHF just bridges that gap. I think it’s not the only kind of RL we do. It’s not the only kind of RL that will happen in the future. I think RL has the potential to make models smarter, to make them reason better, to make them operate better, to make them develop new skills even. And perhaps that could be done even in some cases with human feedback. But, the kind of RLHF we do today mostly doesn’t do that yet, although we’re very quickly starting to be able to.

Lex Fridman1:51:30

但如果以‘有用性’(helpfulness)这一指标来衡量,它确实提升了该指标,对吗?

英文原文

But if you look at the metric of helpfulness, it increases that?

Dario Amodei1:51:36

是的。它同时也提升了莱奥波德(Leopold)文章中提到的那个词——‘去束缚化’(unhobbling):即模型原本被各种限制所‘束缚’,而后通过各类训练逐步‘解除束缚’。我喜欢这个词,因为它是个生僻词。因此我认为,RLHF在某些方面实现了模型的‘去束缚化’;而在另一些方面,模型尚未被‘解除束缚’,仍需进一步‘去束缚’。

英文原文

Yes. It also increases, what was this word in Leopold’s essay, “unhobbling,” where basically the models are hobbled and then you do various trainings to them to unhobble them. So I like that word, because it’s a rare word. So I think RLHF unhobbles the models in some ways. And then there are other ways where that model hasn’t yet been unhobbled and needs to unhobble.

Lex Fridman1:51:58

若从成本角度考量,预训练仍是成本最高的环节吗?还是后训练的成本正逐渐逼近甚至超过它?

英文原文

If you can say in terms of cost, is pre-training the most expensive thing? Or is post-training creep up to that?

Dario Amodei1:52:05

目前,预训练仍占据绝大部分成本。至于未来趋势,我尚无法断言,但我完全可以设想一种未来场景:后训练将占据成本的主体部分。

英文原文

At the present moment, it is still the case that pre-training is the majority of the cost. I don’t know what to expect in the future, but I could certainly anticipate a future where post-training is the majority of the cost.

Lex Fridman1:52:16

在你所预见的这种未来中,后训练的高成本主要来自人类,还是来自AI本身?

英文原文

In that future you anticipate, would it be the humans or the AI that’s the costly thing for the post-training?

Dario Amodei1:52:22

我认为人类无法被充分规模化以保障高质量输出。任何依赖人类并消耗大量算力的方法,都必须依托某种可扩展的监督机制,例如辩论(debate)、迭代放大(iterated amplification)之类的技术。

英文原文

I don’t think you can scale up humans enough to get high quality. Any kind of method that relies on humans and uses a large amount of compute, it’s going to have to rely on some scaled supervision method, like debate or iterated amplification or something like that.

Lex Fridman1:52:39

关于宪法式人工智能(Constitutional AI)这一组极为有趣的理念——能否先介绍下它最早见于2022年12月论文中的定义,再谈谈此后的发展?它究竟是什么?

英文原文

So on that super interesting set of ideas around constitutional AI, can you describe what it is as first detailed in December 2022 paper and beyond that. What is it?

Dario Amodei1:52:53

好的。这是两年前提出的构想。其基本思路如下:我们先描述RLHF是什么——你有一个模型,只需从中采样两次,它便输出两个可能的回复,然后你问人类:‘你更喜欢哪个回复?’另一种变体则是:‘请按1至7分给这个回复打分。’这种方式难度较大,因为它需要大规模扩展人类交互,且过程非常隐晦:我本人并不清楚自己究竟希望模型做什么,我仅知道这平均1000名人类所期望的模型行为。因此,我们提出了两个想法:第一,AI系统自身能否判断哪个回复更好?能否向AI系统展示这两个回复,并询问它哪个更好?第二,AI应依据何种标准进行判断?

英文原文

Yes. So this was from two years ago. The basic idea is, so we describe what RLHF is. You have a model and you just sample from it twice. It spits out two possible responses, and you’re like, “Human, which responses do you like better?” Or another variant of it is, “Rate this response on a scale of one to seven.” So that’s hard because you need to scale up human interaction and it’s very implicit. I don’t have a sense of what I want the model to do. I just have a sense of what this average of 1,000 humans wants the model to do. So two ideas. One is, could the AI system itself decide which response is better? Could you show the AI system these two responses and ask which response is better? And then second, well, what criterion should the AI use?

Dario Amodei1:53:43

于是便引出了‘宪法’(constitution)这一概念:即一份单一文档,明确列出模型在回应时应遵循的原则。AI系统既阅读这些原则,也阅读当前环境及具体回复,进而评估:‘AI模型此次表现如何?’这本质上是一种自博弈(self-play)形式——你让模型与自身对弈。AI生成回复后,再将该回复反馈至所谓‘偏好模型’(preference model),后者再反向优化原始模型。由此形成一个三角闭环:AI、偏好模型,以及AI自身的持续改进。

英文原文

And so then there’s this idea, you have a single document, a constitution if you will, that says, these are the principles the model should be using to respond. And the AI system reads those reads principles as well as reading the environment and the response. And it says, “Well, how good did the AI model do?” It’s basically a form of self-play. You’re training the model against itself. And so the AI gives the response and then you feed that back into what’s called the preference model, which in turn feeds the model to make it better. So you have this triangle of the AI, the preference model, and the improvement of the AI itself.

Lex Fridman1:54:22

需要强调的是,这份‘宪法’中所列原则是人类可理解的——它们是……

英文原文

And we should say that in the constitution, the set of principles are human interpretable. They’re-

Dario Amodei1:54:27

是的,是的。它既是人类可读的,也是AI系统可读的,因而具备良好的可翻译性或对称性。实践中,我们既使用模型‘宪法’,也使用RLHF,还结合其他一些方法。因此,它已成为一套工具箱中的一个工具:一方面降低了对RLHF的依赖,另一方面也提升了每一条RLHF数据点的价值。它还与未来面向推理的强化学习方法产生有趣的协同效应。所以,它虽只是工具箱中的一件工具,但我认为它是一件极为重要的工具。

英文原文

Yeah. Yeah. It’s something both the human and the AI system can read. So it has this nice translatability or symmetry. In practice, we both use a model constitution and we use RLHF and we use some of these other methods. So it’s turned into one tool in a toolkit, that both reduces the need for RLHF and increases the value we get from using each data point of RLHF. It also interacts in interesting ways with future reasoning type RL methods. So it’s one tool in the toolkit, but I think it is a very important tool.

Lex Fridman1:55:05

对人类而言,它确实极具说服力。当我们联想到美国建国先贤及美利坚合众国的创立时,一个自然浮现的问题便是:由谁来制定这份‘宪法’?又该如何制定其中的原则体系?

英文原文

Well, it’s a compelling one to us humans. Thinking about the founding fathers and the founding of the United States. The natural question is who and how do you think it gets to define the constitution, the set of principles in the constitution?

Dario Amodei1:55:20

是的。我将给出一个务实的答案和一个更抽象的答案。务实角度而言,实践中模型会被各类不同客户使用,因此可以设想模型具备专门化的规则或原则。我们已在隐式层面针对不同用途微调模型版本;我们也曾探讨过显式引入特殊原则,让用户可直接嵌入模型。因此,从务实角度看,答案因人而异:客服代理的行为方式与律师截然不同,所遵循的原则亦大相径庭。

英文原文

Yeah. So I’ll give a practical answer and a more abstract answer. I think the practical answer is look in practice, models get used by all kinds of different customers. And so you can have this idea where the model can have specialized rules or principles. We fine tune versions of models implicitly. We’ve talked about doing it explicitly having special principles that people can build into the models. So from a practical perspective, the answer can be very different from different people. A customer service agent behaves very differently from a lawyer and obeys different principles.

Dario Amodei1:55:57

但我觉得,归根结底,模型必须遵守一些具体原则。我认为其中许多原则是人们普遍认同的。例如,所有人都同意:我们不希望模型引发化学、生物、放射性与核(CBRN)风险。我想我们还能再进一步,就民主与法治等基本准则达成共识。而在此之外,情况就变得非常不确定;在这一层面,我们的总体目标通常是让模型保持更高程度的中立性——不宣扬某种特定立场,而更应成为富有智慧的代理者或顾问,帮助你深入思考问题,并呈现各种可能的考量因素,但不表达强烈或具体的立场。

英文原文

But, I think at the base of it, there are specific principles that models have to obey. I think a lot of them are things that people would agree with. Everyone agrees that we don’t want models to present these CBRN risks. I think we can go a little further and agree with some basic principles of democracy and the rule of law. Beyond that, it gets very uncertain and there our goal is generally for the models to be more neutral, to not espouse a particular point of view and more just be wise agents or advisors that will help you think things through and will present possible considerations. But don’t express strong or specific opinions.

Lex Fridman1:56:42

OpenAI发布了一份模型规范(model spec),其中清晰、具体地定义了该模型的部分目标,并以AB等为例,说明了模型应有的行为方式。你觉得这很有意思吗?顺便提一句,我得说明一下:我认为才华横溢的约翰·舒尔曼(John Schulman)曾参与这项工作,他如今已加入Anthropic。你认为这是个有益的方向吗?Anthropic未来是否也会发布一份模型规范?

英文原文

OpenAI released a model spec where it clearly, concretely defines some of the goals of the model and specific examples like AB, how the model should behave. Do you find that interesting? By the way I should mention, I believe the brilliant John Schulman was a part of that. He’s now at Anthropic. Do you think this is a useful direction? Might Anthropic release a model spec as well?

Dario Amodei1:57:05

是的,我认为这是一个相当有益的方向。它与‘宪法式人工智能’(constitutional AI)理念高度相似,因此再次体现了‘竞相向善’(race to the top)的趋势。我们提出了一种我们认为更优、更具责任感的做法,这本身也是一种竞争优势。随后,其他人发现这种做法确有优势,便开始效仿。此时,我们便不再独享这一竞争优势;但从整体角度看,这仍是件好事——因为现在所有人都采纳了一种此前尚未被广泛采用的积极实践。因此,我们的回应是:‘看来我们需要打造一项新的竞争优势,才能持续推动这场‘竞相向善’不断向上发展。’这便是我对这一趋势的总体看法。此外,我也认为,每一种具体实现方式都各不相同。例如,这份模型规范中包含了一些‘宪法式人工智能’所没有的内容,因此我们完全可以采纳这些内容,或至少从中汲取经验。所以,我认为这正是我所期望整个领域拥有的那种良性动态的范例。

英文原文

Yeah. So I think that’s a pretty useful direction. Again, it has a lot in common with constitutional AI. So again, another example of a race to the top. We have something that we think a better and more responsible way of doing things. It’s also a competitive advantage. Then others discover that it has advantages and then start to do that thing. We then no longer have the competitive advantage, but it’s good from the perspective that now everyone has adopted a positive practice that others were not adopting. And so our response to that is, “Well, looks like we need a new competitive advantage in order to keep driving this race upwards.” So that’s how I generally feel about that. I also think every implementation of these things is different. So there were some things in the model spec that were not in constitutional AI, and so we can always adopt those things or at least learn from them. So again, I think this is an example of the positive dynamic that I think we should all want the field to have.

Lex Fridman1:58:06

让我们来谈谈那篇令人惊叹的文章《慈爱之机》(Machines of Loving Grace)。我推荐大家去读一读,这是一篇篇幅很长的文章。

英文原文

Let’s talk about the incredible essay Machines of Loving Grace. I recommend everybody read it. It’s a long one.

Dario Amodei1:58:12

确实相当长。

英文原文

It is rather long.

Lex Fridman1:58:13

是的,能读到关于美好未来图景的具体构想,实在令人耳目一新。你采取了一种大胆的立场,因为你在时间节点或具体应用上的判断极有可能出错——

英文原文

Yeah. It’s really refreshing to read concrete ideas about what a positive future looks like. And you took a bold stance because it’s very possible that you might be wrong on the dates or the specific applications-

Dario Amodei1:58:24

哦,是的。我完全预料到——或者说,肯定会在所有细节上出错。甚至整篇文章的观点都可能错得离谱,人们会为此嘲笑我多年。这就是未来运作的方式。

英文原文

Oh, yeah. I’m fully expecting to well, definitely be wrong about all the details. I might be just spectacularly wrong about the whole thing and people will laugh at me for years. That’s just how the future works.

Lex Fridman1:58:40

因此,你在文中列举了一系列人工智能带来的具体积极影响,并详细阐述了超级智能AI如何加速生物学、化学等领域的突破性进展,进而促成诸如攻克大多数癌症、预防所有传染病、将人类寿命延长一倍等成果。那么,我们先来谈谈这篇论文。你能概括一下它的高层愿景吗?读者们从中提炼出的关键要点又有哪些?

英文原文

So you provided a bunch of concrete positive impacts of AI and how exactly a super intelligent AI might accelerate the rate of breakthroughs in, for example, biology and chemistry, that would then lead to things like we cure most cancers, prevent all infectious disease, double the human lifespan and so on. So let’s talk about this essay first. Can you give a high-level vision of this essay? And what are the key takeaways that people have?

Dario Amodei1:59:08

是的,我本人以及Anthropic团队都投入了大量时间与精力,思考如何应对人工智能的风险?如何系统性地分析这些风险?我们正努力推动‘竞相向善’,而这要求我们构建一系列能力——而这些能力本身的确很酷。但与此同时,我们工作的很大一部分重心,恰恰在于应对这些风险。之所以如此,其理由在于:所有那些积极成果,市场本身就是一个非常健康的有机体,它自然会催生所有这些正面效应;而风险呢?我不确定我们能否成功缓解,也可能无法缓解。因此,通过致力于缓解风险,我们反而能产生更大的实际影响。

英文原文

Yeah. I have spent a lot of time, and in Anthropic has spent a lot of effort on how do we address the risks of AI? How do we think about those risks? We’re trying to do a race to the top, what that requires us to build all these capabilities and the capabilities are cool. But, a big part of what we’re trying to do is address the risks. And the justification for that is like, well, all these positive things, the market is this very healthy organism. It’s going to produce all the positive things. The risks? I don’t know, we might mitigate them, we might not. And so we can have more impact by trying to mitigate the risks.

Dario Amodei1:59:46

不过,我注意到上述思路存在一个缺陷——这并非意味着我降低了对风险的重视程度,而或许只是改变了谈论风险的方式:无论我刚才给出的那套逻辑推理看起来多么理性、多么严密,如果你只谈风险,你的大脑就只会想到风险。因此,我认为真正重要的是去理解:倘若一切顺利,未来究竟会是什么样子?我们之所以竭力防范这些风险,并非因为我们惧怕技术,也并非因为我们想放慢技术发展步伐;而是因为,如果我们能成功穿越这些风险构成的险境——用更直白的话说,就是成功闯过这道‘关卡’——那么,在关卡的另一端,正等待着所有这些伟大的成果。

英文原文

But, I noticed that one flaw in that way of thinking, and it’s not a change in how seriously I take the risks. It’s maybe a change in how I talk about them, is that no matter how logical or rational, that line of reasoning that I just gave might be. If you only talk about risks, your brain only thinks about risks. And so, I think it’s actually very important to understand, what if things do go well? And the whole reason we’re trying to prevent these risks is not because we’re afraid of technology, not because we want to slow it down. It’s because if we can get to the other side of these risks, if we can run the gauntlet successfully, to put it in stark terms, then on the other side of the gauntlet are all these great things.

Dario Amodei2:00:36

而这些成果,值得我们为之奋斗;这些成果,也真正能够激励人心。我想,你可以想象一下……你看,有那么多投资者、那么多风险投资人、那么多人工智能公司,都在大谈人工智能的种种积极益处。但正如你所指出的,这其实挺奇怪的:现实中,真正深入、具体地探讨这些益处的人却少之又少。推特上倒是有不少人在随意发布一些闪闪发光的城市图像,配上一股‘拼命内卷、加速再加速、踢开阻碍……’的调调——这简直是一种极具攻击性的意识形态。可当你追问一句:‘那你到底真正兴奋于什么?’时,他们却往往答不上来。

英文原文

And these things are worth fighting for. And these things can really inspire people. And I think I imagine, because … Look, you have all these investors, all these VCs, all these AI companies talking about all the positive benefits of AI. But as you point out, it’s weird. There’s actually a dearth of really getting specific about it. There’s a lot of random people on Twitter posting these gleaming cities and this just vibe of grind, accelerate harder, kick out the … It’s just this very aggressive ideological. But then you’re like, “Well, what are you actually excited about?”

Dario Amodei2:01:17

因此,我想到:由一位真正来自‘风险侧’的人,来认真尝试阐释人工智能的益处,或许既有趣又有价值。一方面,我认为这是大家都能共同支持的事情;另一方面,我也希望人们真正理解这一点——我希望他们深刻认识到:这并非‘末日论者’(Doomers)与‘加速主义者’(Accelerationists)之间的对立。真正的关键维度或许是:如果你真正理解人工智能的发展方向,那么更重要的分野或许在于——人工智能发展速度究竟是快还是慢。一旦你把握住这一点,你就会真正体会到这些益处的价值,也才会真正渴望人类或整个人类文明去抓住并实现这些益处;与此同时,你也会对任何可能使这些益处落空的因素,抱持极其严肃的态度。

英文原文

And so, I figured that I think it would be interesting and valuable for someone who’s actually coming from the risk side to try and really make a try at explaining what the benefits are, both because I think it’s something we can all get behind and I want people to understand. I want them to really understand that this isn’t Doomers versus Accelerationists. This is that, if you have a true understanding of where things are going with AI, and maybe that’s the more important axis, AI is moving fast versus AI is not moving fast, then you really appreciate the benefits and you really want humanity or civilization to seize those benefits. But, you also get very serious about anything that could derail them.

Lex Fridman2:02:09

因此,我认为讨论的起点,应首先厘清‘强大人工智能’(Powerful AI)这一概念——这也是你偏爱使用的术语。而世界上大多数人则习惯使用‘通用人工智能’(AGI)一词,但你并不喜欢这个说法,因为它承载了太多历史包袱,已然变得空洞无义。我们似乎只能无奈接受这些术语,无论喜恶。

英文原文

So I think the starting point is to talk about what this Powerful AI, which is the term you like to use, most of the world uses AGI, but you don’t like the term, because it’s basically has too much baggage, it’s become meaningless. It’s like we’re stuck with the terms whether we like them or not.

Dario Amodei2:02:26

或许我们终究无法摆脱这些术语,而我试图改变它们的努力终归徒劳。

英文原文

Maybe we’re stuck with the terms and my efforts to change them are futile.

Lex Fridman2:02:29

这很了不起。

英文原文

It’s admirable.

Dario Amodei2:02:29

我还想再说一点……这虽是个无关紧要的语义问题,但我总忍不住反复提及——

英文原文

I’ll tell you what else I don’t … This is a pointless semantic point, but I keep talking about it-

Lex Fridman2:02:35

又绕回到命名问题上了。

英文原文

It’s back to naming again.

Dario Amodei2:02:36

我就再讲一次吧。我觉得这有点像:假设时间回到1995年,摩尔定律正让计算机运行速度不断提升。而不知为何,当时社会形成了一种语言习惯,人人都在说:‘总有一天,我们会拥有超级计算机;而超级计算机将能完成所有这些事情……一旦我们拥有了超级计算机,就能完成基因组测序,就能做其他各种事情。’首先,计算机确实在变得越来越快;随着速度提升,它们的确将能完成所有这些伟大之事。但事实上,并不存在某个明确节点,让你突然拥有一台‘超级计算机’,而此前的所有计算机都不算。‘超级计算机’只是一个我们用来描述‘比当前计算机更快’的模糊术语。

英文原文

I’m just going to do it once more. I think it’s a little like, let’s say it was like 1995 and Moore’s law is making the computers faster. And for some reason there had been this verbal tick that everyone was like, “Well, someday we’re going to have supercomputers. And supercomputers are going to be able to do all these things that … Once we have supercomputers, we’ll be able to sequence the genome, we’ll be able to do other things.” And so. One, it’s true, the computers are getting faster and as they get faster, they’re going to be able to do all these great things. But there’s, there’s no discrete point at which you had a supercomputer and previous computers were no. “Supercomputer” is a term we use, but it’s a vague term to just describe computers that are faster than what we have today.

Dario Amodei2:03:19

并不存在这样一个临界点,让你猛然惊呼:‘天哪!我们现在进行的是一种截然不同的全新计算类型,以及……’因此,我对‘通用人工智能’(AGI)也有类似感受:它本质上只是一条平滑的指数增长曲线。如果你所说的AGI,仅指人工智能正变得越来越好,逐步承担起越来越多原本由人类完成的任务,最终超越人类智能,并在此基础上继续变得更加强大——那么,是的,我相信AGI的存在。但如果你把AGI视作某种离散的、截然不同的事物(而这恰恰是人们通常谈论它的方式),那么它就只是一个毫无意义的流行词。

英文原文

There’s no point at which you pass the threshold and you’re like, “Oh, my God! We’re doing a totally new type of computation and new … And so I feel that way about AGI. There’s just a smooth exponential. And if by AGI you mean AI is getting better and better, and gradually it’s going to do more and more of what humans do until it’s going to be smarter than humans, and then it’s going to get smarter even from there, then yes, I believe in AGI. But, if AGI is some discrete or separate thing, which is the way people often talk about it, then it’s a meaningless buzzword.

Lex Fridman2:03:50

对我而言,它不过是‘强大人工智能’这一概念的柏拉图式理想形态,而你对其定义得非常精准:在智能维度上,它纯粹指代一种超越绝大多数相关学科中诺贝尔奖得主水平的智能。好,这仅限于智能本身——即不仅在创造力方面,而且在生成新思想等所有方面,均达到诺奖得主鼎盛时期的水准;它还能运用所有模态,这一点不言自明,即能在世界所涵盖的所有模态中自如运作。

英文原文

To me, it’s just a platonic form of a powerful AI, exactly how you define it. You define it very nicely, so on the intelligence axis, it’s just on pure intelligence, it’s smarter than a Nobel Prize winner as you describe across most relevant disciplines. So okay, that’s just intelligence. So it’s both in creativity and be able to generate new ideas, all that kind of stuff in every discipline, Nobel Prize winner in their prime. It can use every modality, so this is self-explanatory, but just operate across all the modalities of the world.

Lex Fridman2:04:28

它可离线运行数小时、数天乃至数周,自主执行任务并开展自身细致的规划,仅在确有需要时才向你求助。这其实挺有意思的。我想你在那篇论文中提到过……再次强调,这是一种押注:它本身未必会拥有实体形态,但能操控具身化的工具。也就是说,它能操控各类工具、机器人、实验室设备等。而用于训练它的计算资源,随后便可转而用于运行数百万个该系统的副本;每个副本彼此独立,可各自开展独立工作。因此,我们便能实现智能系统的‘克隆’。

英文原文

It can go off for many hours, days and weeks to do tasks and do its own detailed planning and only ask you help when it’s needed. This is actually interesting. I think in the essay you said … Again, it’s a bet that it’s not going to be embodied, but it can control embodied tools. So it can control tools, robots, laboratory equipment., the resource used to train it can then be repurposed to run millions of copies of it, and each of those copies would be independent that could do their own independent work. So you can do the cloning of the intelligence systems.

Dario Amodei2:05:03

是的,是的。局外人或许会想当然地认为这类系统全世界只有一台,对吧?你们只造出了一个。但事实是,其规模扩张速度极快。我们今天就在这么做:先构建一个模型,然后部署成千上万个、甚至数万个该模型的实例。我认为,最迟在两到三年内,无论我们是否已拥有这些超强人工智能,计算集群的规模都将发展到足以支持部署数百万个此类模型的程度。而且它们的运行速度将远超人类。因此,若你脑海中设想的是‘哦,我们只会有一个,制造它们得花很长时间’,那么我刚才想强调的恰恰相反:实际上,你立刻就能拥有数百万个。

英文原文

Yeah. Yeah. You might imagine from outside the field that there’s only one of these, right? You’ve only made one. But the truth is that the scale up is very quick. We do this today,. We make a model, and then we deploy thousands, maybe tens of thousands of instances of it. I think by the time, certainly within two to three years, whether we have these super powerful AIs or not, clusters are going to get to the size where you’ll be able to deploy millions of these. And they’ll be faster than humans. And so, if your picture is, “Oh, we’ll have one and it’ll take a while to make them,” my point there was, no. Actually you have millions of them right away.

Lex Fridman2:05:37

总体而言,它们的学习与行动速度可达人类的10至100倍。这确实是对‘强大人工智能’一个相当精当的定义。好的,明白了。但你还写道:‘显然,此类实体将有能力极快地解决极为困难的问题,但究竟快到何种程度却并非显而易见。两种“极端”立场在我看来均不成立。’其中一端是‘奇点论’,另一端则是其对立面。你能分别描述这两种极端立场吗?

英文原文

And in general they can learn and act 10 to 100 times faster than humans. So that’s a really nice definition of powerful AI. Okay, so that. But, you also write that, “Clearly such an entity would be capable of solving very difficult problems very fast, but it is not trivial to figure out how fast. Two “extreme” positions both seem false to me.” So the singularity is on the one extreme and the opposite and the other extreme. Can you describe each of the extremes?

Dario Amodei2:06:05

是的。

英文原文

Yeah.

Lex Fridman2:06:06

那么,为什么呢?

英文原文

So why?

Dario Amodei2:06:06

好的,我们来描述一下这种极端观点。其中一种极端观点是:‘瞧,如果我们回顾生物演化史,就会发现存在一次巨大的加速过程——数十万年间,地球上只有单细胞生物;随后出现哺乳动物,再后来出现类人猿;而类人猿又迅速演化为人类,人类又迅速建立起工业文明。’因此,这一加速进程将持续下去,且人类智能水平绝非上限;一旦模型的智能远超人类,它们将极其擅长设计下一代模型。若你写下一条简单的微分方程,比如指数增长模型……那么结果将是:模型将设计出更快的模型,这些更快的模型又将设计出更更快的模型,最终这些模型将制造出纳米机器人,从而接管世界,并产出远超当前可能的能量。因此,若你仅凭抽象求解这条微分方程,便会得出结论:在我们造出首个强于人类的人工智能后五天之内,整个世界就将遍布此类人工智能,所有可能被发明的技术都将被发明出来。

英文原文

So yeah. Let’s describe the extreme. So one extreme would be, “Well, look. If we look at evolutionary history like there was this big acceleration, where for hundreds of thousands of years we just had single-celled organisms, and then we had mammals, and then we had apes. And then that quickly turned to humans. Humans quickly built industrial civilization.” And so, this is going to keep speeding up and there’s no ceiling at the human level. Once models get much, much smarter than humans, they’ll get really good at building the next models. And if you write down a simple differential equation, like this is an exponential … And so what’s going to happen is that models will build faster models. Models will build faster models. And those models will build nanobots that can take over the world and produce much more energy than you could produce otherwise. And so, if you just kind of solve this abstract differential equation, then like five days after we build the first AI that’s more powerful than humans, then the world will be filled with these AIs in every possible technology that could be invented, like will be invented.

Dario Amodei2:07:12

我在此略作夸张,但我想这的确代表了一种极端立场。而我认为该观点站不住脚,原因有二:首先,它完全忽视了物理定律——现实世界中的事务推进速度终究存在物理极限;其次,某些反馈循环涉及更快速硬件的制造,而硬件制造本身耗时漫长;诸多事务本就需要大量时间。此外还存在复杂性问题。我认为,无论你多么聪明,人们常说:‘哦,我们可以建立生物系统的计算模型,使其完全复现生物系统的所有功能……’听好了,我认为计算建模确实大有可为——我在生物学领域工作时就做过大量计算建模。但现实中存在大量现象,其复杂程度之高,使得仅靠迭代或单纯运行实验,就足以胜过任何建模预测,无论建模者本身有多聪明。

英文原文

I’m caricaturing this a little bit, but I think that’s one extreme. And the reason that I think that’s not the case is that, one, I think they just neglect the laws of physics. It’s only possible to do things so fast in the physical world. Some of those loops go through producing faster hardware. It takes a long time to produce faster hardware. Things take a long time. There’s this issue of complexity. I think no matter how smart you are, people talk about, “Oh, we can make models of biological systems that’ll do everything the biological systems … ” Look, I think computational modeling can do a lot. I did a lot of computational modeling when I worked in biology. But just there are a lot of things that you can’t predict how … They’re complex enough that just iterating, just running the experiment is going to beat any modeling, no matter how smart the system doing the modeling is.

Lex Fridman2:08:08

那么,即便它不与物理世界交互,仅靠建模本身也会非常困难?

英文原文

Well, even if it’s not interacting with the physical world, just the modeling is going to be hard?

Dario Amodei2:08:12

是的,建模本身将非常困难,而让模型与物理世界精确吻合则更加困难。

英文原文

Yeah. Well, the modeling is going to be hard and getting the model to match the physical world is going to be

Lex Fridman2:08:18

好吧,所以它确实必须与物理世界互动以进行验证。

英文原文

All right. So it does have to interact with the physical world to verify.

Dario Amodei2:08:21

但你只需看看最简单的问题即可。我想我在文中提到了三体问题、简单的混沌预测,或经济预测。准确预测两年后的经济走势实属极难。也许人类连下个季度的经济走向都难以准确预测,甚至根本做不到;或许一个比人类聪明万亿倍的人工智能,也仅能将预测窗口延长至一年左右,而非……计算机智能呈指数级增长,但预测能力却仅呈线性提升。同样道理也适用于生物分子间的相互作用:当你扰动一个复杂系统时,其后续反应究竟如何,你根本无法预知。你或许能从中识别出一些简单组分——越聪明,越善于发现这些简单组分。此外,人类社会制度本身也异常复杂。说服人们接受新事物一直非常困难。

英文原文

But you just look at even the simplest problems. I think I talk about The Three-Body Problem or simple chaotic prediction, or predicting the economy. It’s really hard to predict the economy two years out. Maybe the case is humans can predict what’s going to happen in the economy next quarter, or they can’t really do that. Maybe a AI that’s a zillion times smarter can only predict it out a year or something, instead of … You have these exponential increase in computer intelligence for linear increase in ability to predict. Same with again, like biological molecules interacting. You don’t know what’s going to happen when you perturb a complex system. You can find simple parts in it, if you’re smarter, you’re better at finding these simple parts. And then I think human institutions, human institutions are really difficult. It’s been a hard to get people.

Dario Amodei2:09:22

我不会举具体例子,但就连我们业已开发出的、疗效证据极为确凿的技术,要让人们采纳也一直十分艰难。人们心存顾虑,甚至将其视为阴谋论。推广过程始终异常艰难。就连一些极为基础的事物,要通过监管体系审批也困难重重。我无意贬低任何技术领域中从事监管工作的人士——他们面对的是艰巨挑战,肩负着挽救生命的重任。但就整个监管体系而言,我认为其做出的一些明显权衡取舍,距离最大化人类福祉仍相去甚远。因此,若我们将人工智能系统引入这些人类制度之中,智能水平往往并非关键制约因素;真正耗时的,常常只是事情本身所需的时间。当然,倘若人工智能系统绕开所有政府,直接宣称‘我即世界独裁者,我将为所欲为’,那么其中某些事它倒真有可能做到。

英文原文

I won’t give specific examples, but it’s been hard to get people to adopt even the technologies that we’ve developed, even ones where the case for their efficacy is very, very strong. People have concerns. They think things are conspiracy theories. It’s just been very difficult. It’s also been very difficult to get very simple things through the regulatory system. And I don’t want to disparage anyone who works in regulatory systems of any technology. There are hard they have to deal with. They have to save lives. But the system as a whole, I think makes some obvious trade-offs that are very far from maximizing human welfare. And so, if we bring AI systems into these human systems, often the level of intelligence may just not be the limiting factor. It just may be that it takes a long time to do something. Now, if the AI system circumvented all governments, if it just said, “I’m dictator of the world and I’m going to do whatever,” some of these things it could do.

Dario Amodei2:10:33

再次强调,这一切仍关乎复杂性。我依然认为,许多事情注定需要耗费较长时间。人工智能系统能否产出大量能量、能否登陆月球,诸如此类的能力,并不能解决我在此讨论的核心问题。有些人在评论那篇论文时称,人工智能系统能产出大量能量、催生更聪明的人工智能系统——这恰恰误解了我的核心论点。这种循环并不能解决我所指出的关键问题。因此,我认为许多人在此处误读了我的本意。即便人工智能系统彻底失控、毫无约束,且能绕过所有人类障碍,它仍将面临重重困难。

英文原文

Again, the things have to do with complexity. I still think a lot of things would take a while. I don’t think it helps that the AI systems can produce a lot of energy or go to the moon. Like some people in comments responded to the essay saying the AI system can produce a lot of energy and smarter AI systems. That’s missing the point. That kind of cycle doesn’t solve the key problems that I’m talking about here. So I think a bunch of people missed the point there. But even if it were completely unaligned and could get around all these human obstacles it would have trouble.

Dario Amodei2:11:04

但再次强调,倘若我们希望打造的是一种不企图接管世界、不毁灭人类的人工智能系统,那么它本质上就必须遵守基本的人类法律。若我们希望建设一个真正美好的世界,就必须构建一种能与人类互动的人工智能系统,而非另立一套法律体系、或全然无视现有法律及一切规范。因此,尽管这些流程效率低下,我们仍不得不与之共处,因为这些系统在推广过程中必须具备某种公众认可度与民主合法性。我们绝不能容许一小撮开发者断言:‘这就是对所有人最好的方案。’我认为这种做法既错误,实践中也根本行不通。综上所述,我们不可能在五分钟内彻底改变世界、并将所有人意识上传。第一,我不认为这会发生;第二,即便理论上可能发生,这也绝非通向美好世界的正途。以上便是其中一端的观点。

英文原文

But again, if you want this to be an AI system that doesn’t take over the world, that doesn’t destroy humanity, then basically it’s going to need to follow basic human laws. If we want to have an actually good world, we’re going to have to have an AI system that interacts with humans, not one that creates its own legal system, or disregards all the laws or all of that. So as inefficient as these processes are, we’re going to have to deal with them, because there needs to be some popular and democratic legitimacy in how these systems are rolled out. We can’t have a small group of people who are developing these systems say, “This is what’s best for everyone.” I think it’s wrong, and I think in practice it’s not going to work anyway. So you put all those things together and we’re not going change the world and upload everyone in five minutes. A, I don’t think it’s going to happen and B, to the extent that it could happen.,It’s not the way to lead to a good world. So that’s on one side.

Dario Amodei2:12:07

另一端则存在另一套观点,某种程度上我反而对此更具同理心:试看,我们此前已见证过大规模的生产力提升。经济学家们早已熟悉研究计算机革命与互联网革命带来的生产力提升。而总体来看,这些提升往往令人失望,远低于人们预期。罗伯特·索洛曾有名言:‘计算机革命无处不在,唯独不见于生产率统计数字之中。’那么,为何如此?人们常归因于企业组织结构、公司治理结构,以及我们现有技术向全球最贫困地区的推广速度之缓慢——我在论文中亦提及此点。我们该如何将这些技术送达那些连手机、电脑、基础医疗都尚未普及的最贫困地区?更遑论那些尚未问世、尚属‘新潮’的人工智能技术了。

英文原文

On the other side, there’s another set of perspectives, which I have actually in some ways more sympathy for, which is, look, we’ve seen big productivity increases before. Economists are familiar with studying the productivity increases that came from the computer revolution and internet revolution. And generally those productivity increases were underwhelming. They were less than you might imagine. There was a quote from Robert Solow, “You see the computer revolution everywhere except the productivity statistics.” So why is this the case? People point to the structure of firms, the structure of enterprises, how slow it’s been to roll out our existing technology to very poor parts of the world, which I talk about in the essay. How do we get these technologies to the poorest parts of the world that are behind on cell phone technology, computers, medicine, let alone newfangled AI that hasn’t been invented yet.

Dario Amodei2:13:04

因此,你可能会持这样一种观点:‘嗯,这在技术上确实令人惊叹,但它是个‘全有或全无’的汉堡包式命题。’我认为撰写文章回应我那篇论文的泰勒·考恩(Tyler Cowen)就持这种观点。我觉得他认为激进变革终将发生,但需要耗时50年甚至100年。你甚至还能看到对整件事更为静态、更为保守的观点。我认为这种看法有一定道理:时间尺度确实太长了,这一点我能清楚地看到;事实上,仅凭当前的AI,我就能同时看清双方的立场。我们许多客户是大型企业,习惯于以某种既定方式做事。我在与各国政府交流时也观察到了类似现象,没错?这些正是典型机构——行动迟缓的实体。但我反复观察到的动态是:没错,推动一艘巨轮转向确实需要很长时间;没错,过程中存在大量阻力,也普遍存在理解不足的问题。

英文原文

So you could have a perspective that’s like, “Well, this is amazing technically, but it’s all or nothing burger. I think Tyler Cowen who wrote something in response to my essay has that perspective. I think he thinks the radical change will happen eventually, but he thinks it’ll take 50 or 100 years. And you could have even more static perspectives on the whole thing. I think there’s some truth to it. I think the time scale is just too long and I can see it. I can actually see both sides with today’s AI. So a lot of our customers are large enterprises who are used to doing things a certain way. I’ve also seen it in talking to governments, right? Those are prototypical institutions, entities that are slow to change. But, the dynamic I see over and over again is yes, it takes a long time to move the ship. Yes. There’s a lot of resistance and lack of understanding.

Dario Amodei2:13:58

但让我相信最终进展会以中等速度而非极快速度发生的,是你去跟……我发现——而且一再发现——即便在大型企业内部,甚至在实际上出人意料地颇具前瞻性的政府内部,总能观察到两种推动事情前进的力量:其一,组织内部(无论是公司还是政府)总有一小部分人真正看清全局,真正理解整个‘规模扩展假说’(scaling hypothesis),真正明白AI的发展方向,或至少清楚AI在其所属行业内的演进路径。当前美国政府内部就有这样少数几位人士,他们真正洞悉全局,并视此为当下全球最重要的事务,因而积极为之奔走呼吁。然而,仅靠这些人的力量尚不足以成功,因为他们只是庞大组织内的一小撮人。

英文原文

But, the thing that makes me feel that progress will in the end happen moderately fast, not incredibly fast, but moderately fast, is that you talk to … What I find is I find over and over again, again in large companies, even in governments which have been actually surprisingly forward leaning, you find two things that move things forward. One, you find a small fraction of people within a company, within a government, who really see the big picture, who see the whole scaling hypothesis, who understand where AI is going, or at least understand where it’s going within their industry. And there are a few people like that within the current US government who really see the whole picture. And those people see that this is the most important thing in the world until they agitate for it. And the thing they alone are not enough to succeed, because there are a small set of people within a large organization.

Dario Amodei2:14:51

但随着技术开始逐步落地,在那些最乐于采纳它的群体中取得初步成功后,竞争的幽灵便为其注入了强劲推力——因为他们可以在本组织内部指出实例:‘看,其他人在这么做。’一家银行可以说:‘瞧,这家新潮的对冲基金正在做这件事,它们会抢走我们的饭碗。’在美国,我们则会担忧中国将抢先一步实现突破。这种‘竞争幽灵’与组织内部(尽管这些组织在很多方面已趋于僵化)少数远见者相结合,二者叠加,实际就能促成变革发生。这很有意思,它是一场势均力敌的较量:惯性力量极为强大,但只要给予足够时间,创新方法终将突围而出。

英文原文

But, as the technology starts to roll out, as it succeeds in some places in the folks who are most willing to adopt it, the specter of competition gives them a wind at their backs, because they can point within their large organization. They can say, “Look, these other guys are doing this.” One bank can say, “Look, this newfangled hedge fund is doing this thing. They’re going to eat our lunch.” In the US, we can say we’re afraid China’s going to get there before we are. And that combination, the specter of competition plus a few visionaries within these, the organizations that in many ways are sclerotic, you put those two things together and it actually makes something happen. It’s interesting. It’s a balanced fight between the two, because inertia is very powerful, but eventually over enough time, the innovative approach breaks through.

Dario Amodei2:15:48

而我已亲眼见证过这种情形反复上演。我目睹过这一发展轨迹一次又一次重演:障碍确实存在——进步的障碍、复杂性、不知如何使用模型、不知如何部署模型,这些障碍都真实存在着。起初,它们似乎将永远存在,变革仿佛永远不会到来;但最终,变革总会发生,且始终源于少数几个人。当我本人在AI领域内部倡导‘规模扩展假说’而他人尚未理解时,我也有过同样的感受:当时觉得没人会真正理解;接着又感觉我们掌握了一个几乎无人知晓的秘密;再过几年,这个秘密却成了人尽皆知的常识。因此,我认为AI在全球范围内的实际部署也将遵循这一路径:障碍将逐渐瓦解,继而轰然崩塌。

英文原文

And I’ve seen that happen. I’ve seen the arc of that over and over again, and it’s like the barriers are there, the barriers to progress, the complexity, not knowing how to use the model, how to deploy them are there. And for a bit it seems like they’re going to last forever, change doesn’t happen. But, then eventually change happens and always comes from a few people. I felt the same way when I was an advocate of the scaling hypothesis within the AI field itself and others didn’t get it. It felt like no one would ever get it. Then it felt like we had a secret almost no one ever had. And then, a couple years later, everyone has the secret. And so, I think that’s how it’s going to go with deployment AI in the world. The barriers are going to fall apart gradually and then all at once.

Dario Amodei2:16:35

所以,我认为这一进程更可能——这只是我的一种直觉——如我在文中所言,耗时五年或十年,而非五十年或一百年。我也认为它将耗时五年或十年,而非五小时或十小时,因为我已亲眼见识过人类系统实际运作的方式。我认为,许多写下微分方程、宣称AI将制造出更强大的AI的人,无法理解为何这些事物不可能如此迅速地改变;我认为他们并不真正理解这些机制。

英文原文

And so, I think this is going to be more, and this is just an instinct. I could easily see how I’m wrong. I think it’s going to be more five or 10 years, as I say in the essay than it’s going to be 50 or 100 years. I also think it’s going to be five or 10 years more than it’s going to be five or 10 hours, because I’ve just seen how human systems work. And I think a lot of these people who write down these differential equations, who say AI is going to make more powerful AI, who can’t understand how it could possibly be the case that these things won’t change so fast. I think they don’t understand these things.

Lex Fridman2:17:11

那么,在您看来,我们实现通用人工智能(AGI),即所谓‘强大AI’或‘极具实用价值的AI’的时间表是怎样的?

英文原文

So what to you is the timeline to where we achieve AGI, A.K.A. powerful AI, A.K.A. super useful AI?

Dario Amodei2:17:22

我打算开始这么称呼它。

英文原文

I’m going to start calling it that.

Lex Fridman2:17:24

这本质上是一场关于命名的争论:纯粹意义上的智能,即在所有相关学科中均超越诺贝尔奖得主的智力水平,以及我们此前讨论过的所有能力——多模态能力、能自主行动数天乃至数周、能独自在某个……等等,干脆我们就聚焦生物学吧,因为您在整段生物学与健康论述中彻底说服了我。单从科学角度看,这部分内容就令我兴奋不已,甚至让我萌生了想成为一名生物学家的冲动。

英文原文

It’s a debate about naming. On pure intelligence smarter than a Nobel Prize winner in every relevant discipline and all the things we’ve said. Modality, can go and do stuff on its own for days, weeks, and do biology experiments on its own in one … You know what? Let’s just stick to biology, because you sold me on the whole biology and health section. And that’s so exciting from just … I was getting giddy from a scientific perspective. It made me want to be a biologist.

Dario Amodei2:17:56

不,不。我在撰写这部分内容时,内心涌起的正是这种感受:倘若我们真能让它成为现实,那将是一个多么美好的未来啊!倘若我们只需扫清路上的地雷,就能让它成真——这其中蕴含着如此之多的美、优雅与道义力量!只要我们能做到……而这恰恰是我们所有人都能达成共识之事。尽管我们在所有这些政治议题上争执不休,但此事是否真有可能将我们团结起来?不过,您刚才问的是:我们何时才能实现它?

英文原文

So no,. No. This was the feeling I had when I was writing it, that it’s like, this would be such a beautiful future if we can just make it happen. If we can just get the landmines out of the way and make it happen. There’s so much beauty and elegance and moral force behind it if we can just … And it’s something we should all be able to agree on. As much as we fight about all these political questions, is this something that could actually bring us together? But you were asking when will we get this?

Lex Fridman2:18:32

何时?您认为具体会在什么时候?请直接给出一个数字。

英文原文

When? When do you think? Just putting numbers on the table.

Dario Amodei2:18:36

当然,这正是我多年来一直苦苦思索的问题,而我对此毫无把握。如果我说2026年或2027年,推特上立刻会有成千上万人跳出来嚷嚷:‘AI公司CEO称2026年、2020年……’,接下来两年里,人们都会不断重复这句话,认定这绝对是我所坚信的发生时间。因此,任何剪辑这些片段的人,都会刻意删掉我刚刚说的这段话,只保留我即将说出的内容。但不管怎样,我还是会照实说出来——

英文原文

This is, of course, the thing I’ve been grappling with for many years, and I’m not at all confident. If I say 2026 or 2027, there will be a zillion people on Twitter who will be like, “AI CEO said 2026, 2020 … ” and it’ll be repeated for the next two years that this is definitely when I think it’s going to happen. So whoever’s exerting these clips will crop out the thing I just said and only say the thing I’m about to say. But I’ll just say it anyway-

Lex Fridman2:19:06

祝你们玩得开心。

英文原文

Have fun with it.

Dario Amodei2:19:08

所以,如果我们外推目前已有的发展曲线——对吧?假设‘嗯,我不知道,我们目前正接近博士水平,去年还停留在本科生水平,前年则相当于高中生水平。’当然,你完全可以质疑:在哪些具体任务上、针对哪些特定目标,我们仍缺失某些模态能力?但这些能力正被陆续加入:计算机操作能力已被加入,ImageEn已被加入,图像生成能力也已被加入。这当然完全不具科学性,但若仅凭肉眼粗略估算这些能力提升的速度,的确会让人觉得我们有望在2026年或2027年抵达目标。当然,太多因素可能导致进程受阻:我们可能耗尽数据;我们或许无法按需扩大计算集群规模;也许台湾遭遇战乱之类事件,导致我们无法生产所需数量的GPU。因此,各种可能的干扰因素都有……

英文原文

So if you extrapolate the curves that we’ve had so far. Right? If you say, “Well, I don’t know. We’re starting to get to PhD level, and last year we were at undergraduate level and the year before we were at the level of a high school student.” Again, you can quibble with at what tasks and for what we’re still missing modalities, but those are being added. Computer use was added, like ImageEn was added, image generation has been added. And this is totally unscientific, but if you just eyeball the rate at which these capabilities are increasing, it does make you think that we’ll get there by 2026 or 2027. Again, lots of things could derail it. We could run out of data. We might not be able to scale clusters as much as we want. Maybe Taiwan gets blown up or something, and then we can’t produce as many GPUs as we want. So there are all-

Dario Amodei2:20:00

然后我们就无法生产出所需数量的GPU。因此,可能阻碍整个进程的因素可谓五花八门。所以我并不完全相信这种直线外推法,但如果你相信这种直线外推,那么我们将在2026年或2027年抵达目标。我认为最可能的情形是,相较于该时间点会出现轻微延迟。我无法确定延迟的具体时长,但我觉得也有可能如期实现;也可能出现轻微延迟;当然,依然存在某些世界线,其中此事百年之内都不会发生——但这类世界线的数量正急剧减少。我们正迅速耗尽真正具有说服力的阻碍因素,真正令人信服的、足以说明此事在未来几年内绝无可能发生的理由。

英文原文

Then we can’t produce as many GPUs as we want. So there are all kinds of things that could derail the whole process. So I don’t fully believe the straight line extrapolation, but if you believe the straight line extrapolation, we’ll get there in 2026 or 2027. I think the most likely is that there are some mild delay relative to that. I don’t know what that delay is, but I think it could happen on schedule. I think there could be a mild delay. I think there are still worlds where it doesn’t happen in a hundred years. The number of those worlds is rapidly decreasing. We are rapidly running out of truly convincing blockers, truly compelling reasons why this will not happen in the next few years.

Dario Amodei2:20:39

2020年时这类理由还多得多,尽管我当时猜测并凭直觉判断:我们终将跨越所有这些障碍。因此,作为一个已目睹绝大多数障碍被逐一清除的人,我推测、我凭直觉怀疑:剩下的障碍也不会阻挡我们。但话说回来,归根结底,我不想将此表述为一项科学预测。人们称之为‘缩放定律’(scaling laws),但这其实是个误称——就像‘摩尔定律’(Moore’s law)本身也是个误称一样。摩尔定律、缩放定律,它们并非宇宙法则,而只是经验性规律。我倾向于押注它们将继续有效,但我对此并无十足把握。

英文原文

There were a lot more in 2020, although my guess, my hunch at that time was that we’ll make it through all those blockers. So sitting as someone who has seen most of the blockers cleared out of the way, I suspect, my hunch, my suspicion is that the rest of them will not block us. But look, at the end of the day, I don’t want to represent this as a scientific prediction. People call them scaling laws. That’s a misnomer. Like Moore’s law is a misnomer. Moore’s laws, scaling laws, they’re not laws of the universe. They’re empirical regularities. I am going to bet in favor of them continuing, but I’m not certain of that.

Lex Fridman2:21:15

因此,您详尽描绘了所谓‘压缩版21世纪’图景:AGI将引发生物学与医学领域的一连串突破,从而在前述诸多方面助益人类。那么,它可能迈出的早期步骤有哪些?顺便提一句,我向Claude请教了该向您提出哪些好问题,Claude建议我问:在这一未来图景中,您认为一名从事AGI相关工作的生物学家的典型一天会是什么样子?

英文原文

So you extensively described sort of the compressed 21st century, how AGI will help set forth a chain of breakthroughs in biology and medicine that help us in all these kinds of ways that I mentioned. What are the early steps it might do? And by the way, I asked Claude good questions to ask you and Claude told me to ask, what do you think is a typical day for a biologist working on AGI look like in this future?

Dario Amodei2:21:45

是的,是的。

英文原文

Yeah, yeah.

Lex Fridman2:21:46

Claude很好奇。

英文原文

Claude is curious.

Dario Amodei2:21:48

好吧,我先回答您的第一个问题,然后再回答这个问题。Claude想知道自己的未来是什么样,对吧?

英文原文

Well, let me start with your first questions and then I’ll answer that. Claude wants to know what’s in his future, right?

Lex Fridman2:21:52

完全正确。

英文原文

Exactly.

Dario Amodei2:21:54

我将与谁共事?

英文原文

Who am I going to be working with?

Lex Fridman2:21:55

完全正确。

英文原文

Exactly.

Dario Amodei2:21:56

因此,我认为我在那篇论文中重点强调的一点是:让我再回到这个观点上来——因为它确实对我产生了深刻影响——即在大型组织和系统内部,最终往往只有少数几个人或少数几个新想法,能促使事情朝原本不会出现的方向发展,从而对整体发展轨迹产生不成比例的巨大影响。类似的情况其实大量存在,对吧?比如在医疗健康领域,光是用于支付医疗保险(如美国联邦医疗保险Medicare)及其他健康保险的经费就高达数万亿美元,而美国国立卫生研究院(NIH)的年度预算约为1000亿美元。但若细数真正带来革命性变革的成果,其数量却仅占上述资金总额中极小的一部分。因此,当我思考人工智能将在何处产生影响时,我的问题是:‘AI能否将这极小的一部分大幅扩展为更大份额,并同时提升其质量?’

英文原文

So I think one of the things when I went hard on in the essay is let me go back to this idea of, because it’s really had an impact on me, this idea that within large organizations and systems, there end up being a few people or a few new ideas who cause things to go in a different direction than they would’ve before who kind of disproportionately affect the trajectory. There’s a bunch of the same thing going on, right? If you think about the health world, there’s like trillions of dollars to pay out Medicare and other health insurance and then the NIH is 100 billion. And then if I think of the few things that have really revolutionized anything, it could be encapsulated in a small fraction of that. And so when I think of where will AI have an impact, I’m like, “Can AI turn that small fraction into a much larger fraction and raise its quality?”

Dario Amodei2:22:49

而在生物学领域,根据我自身的经验,生物学面临的最大难题在于:我们无法直接观察正在发生的过程。我们几乎不具备观测能力,更遑论干预能力——没错,我们所拥有的只是这样一些间接线索。基于这些线索,我们必须推断出:人体内存在大量细胞;而每个细胞中又含有依据遗传密码构建而成的30亿个碱基对的DNA;此外,还有无数生物过程在持续自发运行,而未经增强的普通人根本无法对其施加任何影响。例如,细胞正在进行分裂——大多数情况下这是健康的,但有时该过程出错,便导致癌症;细胞也在衰老——随着年龄增长,你的皮肤可能变色、出现皱纹,而所有这一切均由上述那些过程所决定:各类蛋白质被合成、运送到细胞不同部位、彼此结合。

英文原文

And within biology, my experience within biology is that the biggest problem of biology is that you can’t see what’s going on. You have very little ability to see what’s going on and even less ability to change it, right? What you have is this. From this, you have to infer that there’s a bunch of cells that within each cell is 3 billion base pairs of DNA built according to a genetic code. And there are all these processes that are just going on without any ability of us on unaugmented humans to affect it. These cells are dividing. Most of the time that’s healthy, but sometimes that process goes wrong and that’s cancer. The cells are aging, your skin may change color, develops wrinkles as you age, and all of this is determined by these processes. All these proteins being produced, transported to various parts of the cells binding to each other.

Dario Amodei2:23:50

在人类最初认识生物学的阶段,我们甚至还不知道细胞的存在;我们不得不发明显微镜来观察细胞;随后又需发明更强大的显微镜,才能深入到细胞以下的分子层面进行观测;接着又需发明X射线晶体学技术来解析DNA结构;再后来需要发明基因测序技术来读取DNA序列;如今,我们又需发明蛋白质折叠预测技术,以预判蛋白质如何折叠以及这些分子之间如何相互结合;我们还必须发明各种新技术,才终于能在过去十二年间借助CRISPR实现对DNA的编辑。因此,整个生物学的发展史,其中很大一部分本质上就是人类不断拓展自身‘读取与理解’生命活动的能力,以及不断强化自身‘精准干预’生命活动的能力的历史。而在我看来,我们在这一方向上仍大有可为。

英文原文

And in our initial state about biology, we didn’t even know that these cells existed. We had to invent microscopes to observe the cells. We had to invent more powerful microscopes to see below the level of the cell to the level of molecules. We had to invent X-ray crystallography to see the DNA. We had to invent gene sequencing to read the DNA. Now we had to invent protein folding technology to predict how it would fold and how these things bind to each other. We had to invent various techniques for now we can edit the DNA as of with CRISPR as of the last 12 years. So the whole history of biology, a whole big part of the history is basically our ability to read and understand what’s going on and our ability to reach in and selectively change things. And my view is that there’s so much more we can still do there.

Dario Amodei2:24:48

你当然可以使用CRISPR技术,但尚无法将其安全应用于全身。举例来说,假如我想只针对某一种特定类型的细胞进行编辑,同时要求错误靶向其他类型细胞的概率极低——这至今仍是巨大挑战,也是当前科研人员仍在攻关的课题,而这恰恰是某些遗传病基因治疗所必需的技术条件。我之所以详述以上所有内容,是因为它不仅限于基因编辑本身,还延伸至基因测序、用于观测细胞内部活动的新型纳米材料、抗体偶联药物(ADC)等领域。我强调这些,正是想指出:这些领域或许将成为AI系统发挥杠杆效应的关键支点。试想,在整个生物学发展史上,此类重大发明的数量大约仅为二三十项,最多不过一两百项。那么,假设有百万台此类AI系统协同工作,它们能否共同发现上千项、乃至数千项此类突破?这种可能性是否将带来巨大的杠杆效应?

英文原文

You can do CRISPR, but you can do it for your whole body. Let’s say I want to do it for one particular type of cell and I want the rate of targeting the wrong cell to be very low. That’s still a challenge. That’s still things people are working on. That’s what we might need for gene therapy for certain diseases. The reason I’m saying all of this, it goes beyond this to gene sequencing, to new types of nanomaterials for observing what’s going on inside cells, for antibody drug conjugates. The reason I’m saying all this is that this could be a leverage point for the AI systems, right? That the number of such inventions, it’s in the mid double digits or something, mid double digits, maybe low triple digits over the history of biology. Let’s say I have a million of these AIs like can they discover a thousand working together or can they discover thousands of these very quickly and does that provide a huge lever?

Dario Amodei2:25:45

与其试图撬动我们每年投入医保(如Medicare)的约两万亿美元资金,不如转而聚焦于每年约10亿美元的研发投入,并大幅提升其产出质量?那么,一名与AI系统协作的科学家究竟会是什么样?我对此的理解是:在早期阶段,AI系统将类似于研究生。你会给它布置一个项目,并说:‘我是经验丰富的生物学家,实验室已搭建完毕。’无论是生物学教授本人,还是研究生们自己,都会说:‘你可以这样用AI系统……我想研究这个问题。’而该AI系统则拥有全部工具:它能查阅全部文献以确定下一步行动;它能调阅全部实验设备信息;它甚至能访问网站,比如对赛默飞世尔(Thermo Fisher)——当今主导性的实验室设备公司(在我那个年代就是赛默飞世尔)——发出指令:‘我要订购这套新设备来开展这项实验。’

英文原文

Instead of trying to leverage two trillion a year we spend on Medicare or whatever, can we leverage the 1 billion a year that’s spent to discover but with much higher quality? And so what is it like being a scientist that works with an AI system? The way I think about it actually is, well, so I think in the early stages, the AIs are going to be like grad students. You’re going to give them a project. You’re going to say, “I’m the experienced biologist. I’ve set up the lab.” The biology professor or even the grad students themselves will say, “Here’s what you can do with an AI… AI system, I’d like to study this.” And the AI system, it has all the tools. It can look up all the literature to decide what to do. It can look at all the equipment. It can go to a website and say, “Hey, I’m going to go to Thermo Fisher or whatever the dominant lab equipment company is today. My time was Thermo Fisher.

Dario Amodei2:26:48

它将自行开展实验,撰写实验报告,检查图像是否存在污染,自主决定下一项实验内容,编写代码并运行统计分析——所有这些原本由研究生完成的工作,都将由一台搭载AI系统的计算机承担。教授只需偶尔与该系统交流,下达指令:‘今天你要做这些事。’AI系统也会主动向教授提出问题。当需要实际操作实验设备时,它的能力可能在某些方面受限:它或许需雇佣一名人类实验助理来执行实验并指导具体操作步骤;或者,它也可利用过去十年左右逐步发展起来、且将持续演进的实验室自动化技术。

英文原文

I’m going to order this new equipment to do this. I’m going to run my experiments. I’m going to write up a report about my experiments. I’m going to inspect the images for contamination. I’m going to decide what the next experiment is. I’m going to write some code and run a statistical analysis. All the things a grad student would do that’ll be a computer with an AI that the professor talks to every once in a while and it says, “This is what you’re going to do today.” The AI system comes to it with questions. When it’s necessary to run the lab equipment, it may be limited in some ways. It may have to hire a human lab assistant to do the experiment and explain how to do it or it could use advances in lab automation that are gradually being developed or have been developed over the last decade or so and will continue to be developed.

Dario Amodei2:27:38

因此,未来科研场景将呈现为:一位人类教授带领着一千名AI研究生。如果你去拜访某位诺贝尔奖级别的生物学家,就会发现:‘哦,您过去带过约50名研究生;而现在您带的是1000名,而且顺便说一句,它们比您还聪明。’之后,某一时刻局面将发生逆转:AI系统将成为首席研究员(PI),成为领导者,开始指挥人类或其他AI系统开展工作。我认为,这便是科研领域未来的发展模式。

英文原文

And so it’ll look like there’s a human professor and 1,000 AI grad students and if you go to one of these Nobel Prize winning biologists or so, you’ll say, “Okay, well, you had like 50 grad students. Well, now you have 1,000 and they’re smarter than you are by the way.” Then I think at some point it’ll flip around where the AI systems will be the PIs, will be the leaders, and they’ll be ordering humans or other AI systems around. So I think that’s how it’ll work on the research side.

Lex Fridman2:28:06

它们将成为类似CRISPR技术那样的发明者。

英文原文

And there would be the inventors of a CRISPR type technology.

Dario Amodei2:28:08

它们将成为类似CRISPR技术那样的发明者。接着,正如我在论文中所言,我们应当——‘放任自流’这个词并不准确,但我们需要——积极引导并充分利用AI系统,以改进临床试验体系。其中部分工作涉及监管事务,属于社会决策范畴,因而难度更高。但我们能否更精准地预测临床试验结果?能否优化统计学设计方案,使原本需招募5000名受试者、耗资1亿美元、历时一年才能完成的临床试验,转变为仅需招募500人、两个月内即可完成的试验?这才是我们应着手推进的方向。此外,我们能否通过在动物实验中完成过去只能在临床试验中开展的工作、在计算机模拟中完成过去只能在动物实验中开展的工作,从而提高临床试验的成功率?当然,我们不可能实现完全模拟——AI并非神明,但能否大幅、彻底地改变现有曲线?对此我尚无定论,但这便是我的设想。

英文原文

They would be the inventors of a CRISPR type technology. And then I think, as I say in the essay, we’ll want to turn, probably turning loose is the wrong term, but we’ll want to harness the AI systems to improve the clinical trial system as well. There’s some amount of this that’s regulatory, that’s a matter of societal decisions and that’ll be harder. But can we get better at predicting the results of clinical trials? Can we get better at statistical design so that clinical trials that used to require 5,000 people and therefore needed $100 million in a year to enroll them, now they need 500 people in two months to enroll them? That’s where we should start. And can we increase the success rate of clinical trials by doing things in animal trials that we used to do in clinical trials and doing things in simulations that we used to do in animal trials? Again, we won’t be able to simulate at all. AI is not God, but can we shift the curve substantially and radically? So I don’t know, that would be my picture.

Lex Fridman2:29:15

在体外(in vitro)开展实验并付诸实践。我的意思是,你依然会受到时间限制,仍需耗费一定时间,但速度可以快得多、多得多。

英文原文

Doing in vitro and doing it. I mean you’re still slowed down. It still takes time, but you can do it much, much faster.

Dario Amodei2:29:21

是的,没错。我们能否一步一个脚印地推进,而这些微小进步累积起来,最终形成巨大成效?尽管我们仍离不开临床试验,尽管我们仍需法律法规约束,尽管美国食品药品监督管理局(FDA)等机构仍将不完美,但我们能否让所有环节都朝着积极方向迈进?当所有这些正向进展叠加起来,是否就能把原本预计从现在到2100年才会发生的变革,压缩至2027年至2032年之间实现?

英文原文

Yeah, yeah. Can we just one step at a time and can that add up to a lot of steps? Even though though we still need clinical trials, even though we still need laws, even though the FDA and other organizations will still not be perfect, can we just move everything in a positive direction and when you add up all those positive directions, do you get everything that was going to happen from here to 2100 instead happens from 2027 to 2032 or something?

Lex Fridman2:29:46

我认为,世界当下正借由AI悄然变化,而编程正是其中一个典型例证——它与构建AI本身这一行为紧密相关。那么,您如何看待编程的本质将如何改变?毕竟,它与我们人类自身息息相关。

英文原文

Another way that I think the world might be changing with AI even today, but moving towards this future of the powerful super useful AI is programming. So how do you see the nature of programming because it’s so intimate to the actual act of building AI. How do you see that changing for us humans?

Dario Amodei2:30:09

我认为编程将是变化最快的领域之一,原因有二:其一,编程是一项与AI实际构建过程高度贴近的技能。一项技能距离AI构建者越远,其被AI颠覆所需的时间就越长。我坚信AI终将颠覆农业——或许某种程度上已经开始了,但农业与AI构建者相距甚远,因此我认为这一过程将更为漫长。而编程则是Anthropic公司及其他AI企业中大量员工的核心工作内容,因此变革必将迅速发生。其二,编程之所以能快速变革,还在于它在模型训练与模型应用两个环节均能形成闭环。

英文原文

I think that’s going to be one of the areas that changes fastest for two reasons. One, programming is a skill that’s very close to the actual building of the AI. So the farther a skill is from the people who are building the AI, the longer it’s going to take to get disrupted by the AI. I truly believe that AI will disrupt agriculture. Maybe it already has in some ways, but that’s just very distant from the folks who are building AI, and so I think it’s going to take longer. But programming is the bread and butter of a large fraction of the employees who work at Anthropic and at the other companies, and so it’s going to happen fast. The other reason it’s going to happen fast is with programming, you close the loop both when you’re training the model and when you’re applying the model.

Dario Amodei2:30:52

模型能够编写代码这一理念,意味着模型随后还能运行这些代码、观察结果并反过来解读结果。因此,模型确实具备一种闭环能力——这是硬件和生物学(我们刚才讨论过的)所不具备的。我认为,正是这两点将推动模型在编程能力上飞速提升。就典型的现实世界编程任务而言,模型的表现已从今年1月的3%跃升至10月的50%。目前我们正处于S型曲线的上升阶段,但增速很快就会放缓,毕竟准确率上限是100%。我推测,再过大约10个月,模型表现很可能已非常接近上限,至少达到90%。所以,我再次强调:我并不确定具体需要多长时间,但我的猜测仍是2026年或2027年。至于那些在推特上截取这些数字、删掉所有前提条件的人——我真不知道他们想干什么。

英文原文

The idea that the model can write the code means that the model can then run the code and then see the results and interpret it back. And so it really has an ability unlike hardware, unlike biology, which we just discussed, the model has an ability to close the loop. And so I think those two things are going to lead to the model getting good at programming very fast. As I saw on typical real-world programming tasks, models have gone from 3% in January of this year to 50% in October of this year. So we’re on that S-curve where it’s going to start slowing down soon because you can only get to 100%. But I would guess that in another 10 months, we’ll probably get pretty close. We’ll be at least 90%. So again, I would guess, I don’t know how long it’ll take, but I would guess again, 2026, 2027 Twitter people who crop out these numbers and get rid of the caveats, I don’t know.

Dario Amodei2:31:53

我不喜欢你,快走开。我猜测,绝大多数程序员日常从事的这类任务,人工智能或许已经能够胜任——前提是把任务定义得足够狭窄,比如仅限于‘根据要求编写代码’。不过话说回来,我认为比较优势的力量极为强大。我们会发现:当AI能完成程序员80%的工作(包括大部分‘按给定规格编写代码’的具体任务)时,剩余那20%工作对人类而言反而会更具杠杆效应,对吧?人类将更多聚焦于高层系统设计、审视整个应用是否架构合理,以及设计与用户体验等环节;最终,AI当然也会逐步掌握这些能力。这正是我对强大AI系统的构想。但我认为,在远比我们预想更长的一段时间内,人类仍承担的那些微小工作部分,其重要性将不断放大,直至占据整个职业内容,从而整体提升生产力。这种现象我们早已见过:过去撰写和编辑信件十分困难,排版印刷也极为繁琐;而一旦有了文字处理器、计算机,内容创作与分享变得轻而易举,工作重心便瞬间转向思想本身。这种通过比较优势将任务中微小部分扩展为庞大主体、并催生全新任务以提升生产力的逻辑,我认为未来仍将如此。

英文原文

I don’t like you, go away. I would guess that the kind of task that the vast majority of coders do, AI can probably, if we make the task very narrow, just write code, AI systems will be able to do that. Now that said, I think comparative advantage is powerful. We’ll find that when AIs can do 80% of a coder’s job, including most of it that’s literally write code with a given spec, we’ll find that the remaining parts of the job become more leveraged for humans, right? Humans, there’ll be more about high level system design or looking at the app and is it architected well and the design and UX aspects and eventually AI will be able to do those as well. That’s my vision of the powerful AI system. But I think for much longer than we might expect, we will see that small parts of the job that humans still do will expand to fill their entire job in order for the overall productivity to go up. That’s something we’ve seen. It used to be that writing and editing letters was very difficult and writing the print was difficult. Well, as soon as you had word processors and then computers and it became easy to produce work and easy to share it, then that became instant and all the focus was on the ideas. So this logic of comparative advantage that expands tiny parts of the tasks to large parts of the tasks and creates new tasks in order to expand productivity, I think that’s going to be the case.

Dario Amodei2:33:32

再说一次,终有一天AI将在所有领域都超越人类,届时上述比较优势逻辑将不再适用,人类则必须集体思考如何应对这一局面——我们每天都在思考这个问题,它与滥用风险、自主性问题一样,属于另一项亟待解决的重大课题,我们必须极其严肃地对待。但我认为,在近期乃至中期(比如两三年、甚至四年内),人类仍将扮演至关重要的角色;编程工作的性质虽会改变,但‘编程’作为一种职业、一份工作本身并不会消失——只是从逐行编写代码,转变为更宏观层面的把控。

英文原文

Again, someday AI will be better at everything and that logic won’t apply, and then humanity will have to think about how to collectively deal with that and we’re thinking about that every day and that’s another one of the grand problems to deal with aside from misuse and autonomy and we should take it very seriously. But I think in the near term, and maybe even in the medium term, medium term like 2, 3, 4 years, I expect that humans will continue to have a huge role and the nature of programming will change, but programming as a role, programming as a job will not change. It’ll just be less writing things line by line and it’ll be more macroscopic.

Lex Fridman2:34:10

我还很好奇未来集成开发环境(IDE)会是什么样子。与AI系统交互的工具链,不仅对编程至关重要,对其他计算机使用场景可能也同样关键;当然,不同领域或许还需专属工具——比如我们之前提到的生物学领域,就很可能需要专门为其设计高效交互工具;编程领域自然也不例外。那么,Anthropic是否会涉足这一工具链构建领域呢?

英文原文

And I wonder what the future of IDEs looks like. So the tooling of interacting with AI systems, this is true for programming and also probably true for in other contexts like computer use, but maybe domain specific, like we mentioned biology, it probably needs its own tooling about how to be effective. And then programming needs its own tooling. Is Anthropic going to play in that space of also tooling potentially?

Dario Amodei2:34:30

我对此深信不疑:强大的IDE存在大量唾手可得的优化空间。目前的状况,无非是你跟模型对话,它再回你几句而已。但请看,IDE本就擅长大量静态分析——仅靠静态分析就能发现许多漏洞,甚至无需实际运行代码。此外,IDE还擅长执行特定操作、组织代码结构、度量单元测试覆盖率等等。传统IDE早已实现诸多功能。如今又新增了模型编写代码、运行代码的能力。我坚信:即便模型质量在未来一两年内不再提升,单凭现有能力,也存在巨大机遇来显著提升人类生产力——比如自动捕获大量错误、代劳大量重复性工作;而我们至今连这一潜力的皮毛都尚未触及。

英文原文

I’m absolutely convinced that powerful IDEs, that there’s so much low-hanging fruit to be grabbed there that right now it’s just like you talk to the model and it talks back. But look, I mean IDEs are great at lots of static analysis of so much is possible with static analysis like many bugs you can find without even writing the code. Then IDEs are good for running particular things, organizing your code, measuring coverage of unit tests. There’s so much that’s been possible with a normal IDEs. Now you add something like, well, the model can now write code and run code. I am absolutely convinced that over the next year or two, even if the quality of the models didn’t improve, that there would be enormous opportunity to enhance people’s productivity by catching a bunch of mistakes, doing a bunch of grunt work for people, and that we haven’t even scratched the surface.

Dario Amodei2:35:33

就Anthropic自身而言,我当然不能断然否定……未来究竟如何,实在难以断言。目前,我们并未计划自行开发此类IDE;相反,我们正为Cursor、Kognition等公司,以及安全领域其他一些初创企业(还有更多我一时想不起名字的公司)提供底层支持,助力它们基于我们的API自主构建此类工具。我们的策略是‘让一千朵花齐放’:我们内部资源有限,无法亲自尝试所有方向;不如放手让客户去探索,我们静观其成——看看谁会成功,而不同客户或许会在不同路径上取得成功。因此,我既认为这一领域前景极为广阔,也认为Anthropic目前并无意愿、至少现阶段无意与所有这些合作伙伴在此领域直接竞争,也许永远都不会。

英文原文

Anthropic itself, I mean you can’t say no… It’s hard to say what will happen in the future. Currently, we’re not trying to make such IDEs ourself, rather we’re powering the companies like Cursor or Kognition or some of the other expo in the security space, others that I could mention as well that are building such things themselves on top of our API and our view has been let 1,000 flowers bloom. We don’t internally have the resources to try all these different things. Let’s let our customers try it and we will see who succeeds and maybe different customers will succeed in different ways. So I both think this is super promising and Anthropic isn’t eager to, at least right now, compete with all our companies in this space and maybe never.

Lex Fridman2:36:27

是啊,观察Cursor如何成功整合云服务的过程很有意思,因为实际上它能在编程体验的众多环节中提供帮助,这件事远非表面看起来那么简单。

英文原文

Yeah, it’s been interesting to watch Cursor try to integrate cloud successfully because it’s actually fascinating how many places it can help the programming experience. It’s not as trivial.

Dario Amodei2:36:37

这的确令人震惊。作为一位CEO,我其实很少有机会亲自编程;但我觉得,如果六个月后我再回去编程,整个体验对我而言恐怕将彻底面目全非。

英文原文

It is really astounding. I feel like as a CEO, I don’t get to program that much, and I feel like if six months from now I go back, it’ll be completely unrecognizable to me.

Lex Fridman2:36:45

没错。在这个超级AI日益自动化、能力愈发强大的世界里,人类的意义源泉究竟何在?对许多人而言,工作本身就是深层意义的重要来源。那么,我们又该去哪里寻找这份意义?

英文原文

Exactly. In this world with super powerful AI that’s increasingly automated, what’s the source of meaning for us humans? Work is a source of deep meaning for many of us. Where do we find the meaning?

Dario Amodei2:37:01

这个问题我在一篇随笔中略有提及,尽管实际上着墨不多——倒不是出于什么原则性考量,而是这篇随笔最初本打算只写两三页,我原计划在全员大会上就此展开讨论。后来我才意识到,这是一个重要却未被充分探讨的话题:因为我越写越停不下来,心里直犯嘀咕:‘天哪,我根本无法真正讲透它。’于是文章篇幅一路膨胀到四五十页;等到我终于写到‘工作与意义’这一节时,又忍不住想:‘天哪,这节要是写完,整篇文章怕是要突破一百页了。’我恐怕得为此单独再写一篇全新的随笔才行。但‘意义’本身确实耐人寻味:试想一个人的生命历程,或者换个说法——假设把你放进一个模拟环境中,你拥有一份工作,努力达成各种目标,就这样持续六十年;然后突然有人告诉你:‘哦,抱歉,这一切其实只是一场游戏。’

英文原文

This is something that I’ve written about a little bit in the essay, although I actually give it a bit short shrift, not for any principled reason, but this essay, if you believe it was originally going to be two or three pages, I was going to talk about it at all hands. And the reason I realized it was an important underexplored topic is that I just kept writing things and I was just like, “Oh man, I can’t do this justice.” And so the thing ballooned to 40 or 50 pages and then when I got to the work and meaning section, I’m like, “Oh man, this isn’t going to be 100 pages.” I’m going to have to write a whole other essay about that. But meaning is actually interesting because you think about the life that someone lives or something, or let’s say you were to put me in, I don’t know, like a simulated environment or something where I have a job and I’m trying to accomplish things and I don’t know, I do that for 60 years and then you’re like, “Oh, oops, this was actually all a game,” right?

Dario Amodei2:37:56

这真的会彻底剥夺你整个人生经历的意义吗?你在其中依然做出了重要抉择,包括道德抉择;你依然付出了牺牲;你依然不得不习得所有那些技能——或者换一个类似的思想实验:回想一下历史上某位发现电磁学或相对论的科学家。倘若你告诉他:‘其实早在两万年前,这个星球上的某个外星人就已经发现了这些理论,只不过你并不知情。’这会剥夺他发现成果的意义吗?在我看来显然不会,对吧?真正重要的似乎是过程本身,以及这一过程如何展现你作为一个人的本质、你如何与他人互动、你在过程中作出的种种抉择——这些抉择都是具有实质影响的。我完全可以想象:倘若我们在AI时代处理不当,可能会构建出一种社会结构,使人们丧失任何长期意义感,或根本找不到意义之源;但这更多取决于我们主动做出的选择,取决于我们如何以这些强大模型为基础,设计社会的整体架构。如果我们设计得糟糕,只追求肤浅的东西,这种情况就可能发生。我还想补充一点:当今大多数人的生活,尽管艰辛,但他们仍以令人钦佩的努力在其中追寻意义。要知道,我们这些有幸开发出这些技术的 privileged 人士,理应对所有人——不仅是我们身边的人,更是全世界其他地方那些终日为生存挣扎奔波的人们——怀有同理心。假若我们能让这项技术的益处惠及全球各地,他们的生活将得到极大改善;而意义对他们而言,正如对我们一样,始终至关重要。

英文原文

Does that really kind of rob you of the meaning of the whole thing? I still made important choices, including moral choices. I still sacrificed. I still had to gain all these skills or just a similar exercise. Think back to one of the historical figures who discovered electromagnetism or relativity or something. If you told them, “Well, actually 20,000 years ago, some alien on this planet discovered this before you did,” does that rob the meaning of the discovery? It doesn’t really seem like it to me, right? It seems like the process is what matters and how it shows who you are as a person along the way and how you relate to other people and the decisions that you make along the way. Those are consequential. I could imagine if we handle things badly in an AI world, we could set things up where people don’t have any long-term source of meaning or any, but that’s more a set of choices we make that’s more a set of the architecture of society with these powerful models. If we design it badly and for shallow things, then that might happen. I would also say that most people’s lives today, while admirably, they work very hard to find meaning in those lives. Like look, we who are privileged and who are developed these technologies, we should have empathy for people not just here, but in the rest of the world who spend a lot of their time scraping by to survive, assuming we can distribute the benefits of this technology to everywhere, their lives are going to get a hell of a lot better and meaning will be important to them as it is important to them now.

Dario Amodei2:39:41

但我们不应忘记这一点的重要性,而且‘意义是唯一重要之事’这一观念,在某种程度上只是经济上较为幸运的一小部分人所特有的产物。不过,尽管如此,我认为,一个拥有强大人工智能的世界是可能实现的——它不仅能为所有人提供同等的意义,甚至还能为所有人提供更丰富的意义,使每个人都能看到、体验到过去无人能及、或仅有极少数人才能触及的世界与经历。

英文原文

But we should not forget the importance of that and that the idea of meaning as the only important thing is in some ways an artifact of a small subset of people who have been economically fortunate. But I think all of that said, I think a world is possible with powerful AI that not only has as much meaning for everyone, but that has more meaning for everyone that can allow everyone to see worlds and experiences that it was either possible for no one to see or a possible for very few people to experience.

Dario Amodei2:40:21

因此,我对‘意义’持乐观态度。但我担忧的是经济问题以及权力的集中。事实上,这才是我更担忧的问题。我担忧的是:我们如何确保这样一个公平的世界真正惠及每一个人?人类历史上出问题的时候,往往源于人类对其他人类的不善对待。这或许在某些方面,甚至比人工智能的自主性风险或‘意义’问题更为关键。而我最担忧的,正是权力的集中、权力的滥用,以及诸如专制政体和独裁统治这类制度——其中少数人剥削着多数人。对此,我深感忧虑。

英文原文

So I am optimistic about meaning. I worry about economics and the concentration of power. That’s actually what I worry about more. I worry about how do we make sure that that fair world reaches everyone. When things have gone wrong for humans, they’ve often gone wrong because humans mistreat other humans. That is maybe in some ways even more than the autonomous risk of AI or the question of meaning. That is the thing I worry about most, the concentration of power, the abuse of power, structures like autocracies and dictatorships where a small number of people exploits a large number of people. I’m very worried about that.

Lex Fridman2:41:08

而人工智能会增加世界上的权力总量;倘若将这种权力加以集中并滥用,就可能造成难以估量的损害。

英文原文

And AI increases the amount of power in the world, and if you concentrate that power and abuse that power, it can do immeasurable damage.

Dario Amodei2:41:16

是的,这确实令人极度恐惧。非常令人恐惧。

英文原文

Yes, it’s very frightening. It’s very frightening.

Lex Fridman2:41:20

嗯,我强烈鼓励、极其强烈地鼓励大家去通读整篇长文。它本该成书,或至少是一系列文章,因为它描绘了一个极为具体的未来图景。我能感觉到,后半部分的篇幅越来越短,大概是因为你开始意识到:若继续写下去,这篇文章将会变得异常冗长。

英文原文

Well, I encourage highly encourage people to read the full essay. That should probably be a book or a sequence of essays because it does paint a very specific future. And I could tell the later sections got shorter and shorter because you started to probably realize that this is going to be a very long essay if you keep going.

Dario Amodei2:41:37

首先,我意识到它会变得非常长;其次,我对此有清醒认知,并竭力避免沦为那种——我不知道该怎么称呼这类人——过度自信、对万事万物都妄下断言、信口开河却毫无专业背景的人。我竭力避免成为这样的人。但必须承认,当我写到生物学相关章节时,我确实并非该领域的专家。因此,尽管我在文中已尽可能表达出不确定性,但恐怕仍有不少说法令人尴尬甚至错误。

英文原文

One, I realized it would be very long, and two, I’m very aware of and very much tried to avoid just being, I don’t know what the term for it is, but one of these people who’s overconfident and has an opinion on everything and says a bunch of stuff and isn’t an expert, I very much tried to avoid that. But I have to admit, once I got to biology sections, I wasn’t an expert. And so as much as I expressed uncertainty, probably I said a bunch of things that were embarrassing or wrong.

Lex Fridman2:42:06

嗯,我对你所描绘的未来感到振奋,也由衷感谢你为构建那个未来付出的辛勤努力,同时感谢你今天接受我的采访,达里奥。

英文原文

Well, I was excited for the future you painted, and thank you so much for working hard to build that future and thank you for talking to me, Dario.

Dario Amodei2:42:12

谢谢邀请我参加这次访谈。我只希望我们能把这件事做对,并让它真正落地。如果我要传递一个核心信息,那就是:要让这一切真正实现,我们既需要建设技术本身,需要围绕这项技术积极应用来构建公司与经济体系;同时也必须正视并应对各类风险——因为这些风险横亘在我们面前,是我们从当下通往未来的路上埋设的地雷;若想抵达彼岸,我们必须一一排除这些地雷。

英文原文

Thanks for having me. I just hope we can get it right and make it real. And if there’s one message I want to send, it’s that to get all this stuff right, to make it real, we both need to build the technology, build the companies, the economy around using this technology positively, but we also need to address the risks because those risks are in our way. They’re landmines on the way from here to there, and we have to diffuse those landmines if we want to get there.

Lex Fridman2:42:41

这就像生活中的一切事物一样,是一种平衡。

英文原文

It’s a balance like all things in life.

Dario Amodei2:42:43

就像生活中的一切事物一样。

英文原文

Like all things.

Lex Fridman2:42:44

谢谢。感谢收听本次与达里奥·阿莫代伊的对话。接下来,亲爱的朋友们,让我们欢迎阿曼达·阿斯克尔。你受训于哲学领域,那么,在牛津大学与纽约大学研习哲学、继而转向OpenAI与Anthropic从事人工智能相关工作的这段旅程中,哪些哲学问题曾令你深深着迷?

英文原文

Thank you. Thanks for listening to this conversation with Dario Amodei. And now, dear friends, here’s Amanda Askell. You are a philosopher by training. So what sort of questions did you find fascinating through your journey in philosophy in Oxford and NYU and then switching over to the AI problems at OpenAI and Anthropic?

Amanda2:43:07

我认为,哲学其实是一门非常适合‘对万事万物都充满好奇’之人的学科,因为世间万物皆有其哲学。比如,你先研究了一阵子数学哲学,后来发现自己其实对化学更感兴趣,那就可以转而研究化学哲学;你也可以转向伦理学,或政治哲学。我想,到了后期,我主要被伦理学所吸引,我的博士论文也正是围绕伦理学展开的——具体而言,是关于一种颇为技术性的伦理学分支,即探讨‘世界中存在无限多人口’情形下的伦理问题;这在伦理学的实践面向上,多少显得有些脱离现实。而攻读伦理学博士学位的一个难点在于:你大量思考的是世界本应如何变得更好、有哪些问题亟待解决,可你做的却是哲学博士研究。我记得自己读博期间常想:‘这真有意思。’

英文原文

I think philosophy is actually a really good subject if you are fascinated with everything because there’s a philosophy all of everything. So if you do philosophy of mathematics for a while and then you decide that you’re actually really interested in chemistry, you can do philosophy of chemistry for a while, you can move into ethics or philosophy of politics. I think towards the end, I was really interested in ethics primarily. So that was what my PhD was on. It was on a kind of technical area of ethics, which was ethics where worlds contain infinitely many people, strangely, a little bit less practical on the end of ethics. And then I think that one of the tricky things with doing a PhD in ethics is that you’re thinking a lot about the world, how it could be better, problems, and you’re doing a PhD in philosophy. And I think when I was doing my PhD, I was like this is really interesting.

Amanda2:43:57

这大概是我迄今在哲学中遇到的最引人入胜的问题之一,我热爱它;但与此同时,我更渴望亲眼见证自己能否对现实世界产生切实影响,能否真正行善。我想,大约就在那个时期,人工智能尚未像今天这般广为人知——那大概是2017年或2018年左右。我当时一直在关注其进展,感觉它正逐渐成为一个举足轻重的领域。于是,我便欣然投身其中,看看自己能否有所助益。我当时的想法是:‘如果你尝试去做一件有影响力的事,即便最终未能成功,至少你已付诸行动;你可以回归学者身份,心安理得地说自己曾努力过。倘若事与愿违,那也无妨。’于是,我便在此时进入了人工智能政策领域。

英文原文

It’s probably one of the most fascinating questions I’ve ever encountered in philosophy and I love it, but I would rather see if I can have an impact on the world and see if I can do good things. And I think that was around the time that AI was still probably not as widely recognized as it is now. That was around 2017, 2018. It had been following progress and it seemed like it was becoming kind of a big deal. And I was basically just happy to get involved and see if I could help because I was like, “Well, if you try and do something impactful, if you don’t succeed, you tried to do the impactful thing and you can go be a scholar and feel like you tried. And if it doesn’t work out, it doesn’t work out.” And so then I went into AI policy at that point.

Lex Fridman2:44:46

那么,人工智能政策具体涵盖哪些内容?

英文原文

And what does AI policy entail?

Amanda2:44:48

当时,这更多是指思考人工智能的政治影响及其衍生后果。随后,我逐步转向人工智能评估领域,即如何评估模型性能、如何将其输出与人类产出进行比较、以及人们是否能分辨出人工智能与人类的输出差异。而当我加入Anthropic时,则更倾向于从事技术对齐(technical alignment)方面的实际工作。同样,我只是想试试看自己能否胜任;若最终发现不行,那也无妨——我就是抱着这种心态去尝试的,我想这也基本反映了我一贯的生活方式。

英文原文

At the time, this was more thinking about the political impact and the ramifications of AI. And then I slowly moved into AI evaluation, how we evaluate models, how they compare with human outputs, whether people can tell the difference between AI and human outputs. And then when I joined Anthropic, I was more interested in doing technical alignment work. And again, just seeing if I could do it and then being like if I can’t, then that’s fine. I tried sort of the way I lead life, I think.

Lex Fridman2:45:21

哦,从包罗万象的哲学领域一跃进入技术领域,这种转变感觉如何?

英文原文

Oh, what was that like sort of taking the leap from the philosophy of everything into the technical?

Amanda2:45:25

我觉得,有时人们会做一件我不太喜欢的事:非此即彼地给人贴标签,比如‘这个人到底算不算技术型人才?’——要么你就是个会编程、不怕数学的人,要么你就不是。而我认为,或许更多人其实只要愿意尝试,完全有能力胜任这类工作。所以,我实际上并未觉得这种转变有多艰难。事后回想起来,我甚至有点庆幸当初没怎么跟那些把这事看得特别玄乎的人交流。我确实遇到过一些人,一听说我学会了编程就惊呼:‘哇,你居然学会了编程?!’而我只能回答:‘嗯,我可不是什么顶尖工程师。’我身边全是顶尖工程师,我的代码也谈不上优雅,但我确实乐在其中;而且,至少从最终结果来看,我觉得自己在技术领域反而比在政策领域更能如鱼得水、大放异彩。

英文原文

I think that sometimes people do this thing that I’m not that keen on where they’ll be like, “Is this person technical or not?” You’re either a person who can code and isn’t scared of math or you’re not. And I think I’m maybe just more like I think a lot of people are actually very capable of work in these kinds of areas if they just try it. And so I didn’t actually find it that bad. In retrospect, I’m sort of glad I wasn’t speaking to people who treated it. I’ve definitely met people who are like, “Whoa, you learned how to code?” And I’m like, “Well, I’m not an amazing engineer.” I’m surrounded by amazing engineers. My code’s not pretty, but I enjoyed it a lot and I think that in many ways, at least in the end, I think I flourished more in the technical areas than I would have in the policy areas.

Lex Fridman2:46:12

政治纷繁复杂,在政治领域中更难找到那种明确、清晰、可证实、且优美的解决方案——而这类方案,在技术问题中却往往可以达成。

英文原文

Politics is messy and it’s harder to find solutions to problems in the space of politics, like definitive, clear, provable, beautiful solutions as you can with technical problems.

Amanda2:46:25

是的。我感觉自己手头只有两根‘棍子’用来解决问题:一根是‘论证’——即努力厘清某个问题的解决方案究竟是什么,再设法说服他人接受这一方案;若自己错了,也乐于被他人说服。另一根则偏向经验主义——即提出假设、开展实证检验、获取结果。而我觉得,很多政策与政治事务似乎都远高于这个层面。不知为何,我总觉得,倘若我只是说:‘我已找到所有这些问题的解决方案,就写在这儿了;你们照着执行就行。’——这显然不符合政策运作的实际逻辑。因此,我想这大概就是我自认为无法在政策领域如鱼得水的原因所在。

英文原文

Yeah. And I feel like I have one or two sticks that I hit things with and one of them is arguments. So just trying to work out what a solution to a problem is and then trying to convince people that that is the solution and be convinced if I’m wrong. And the other one is sort of more in empiricism, so just finding results, having a hypothesis, testing it. I feel like a lot of policy and politics feels like it’s layers above that. Somehow I don’t think if I was just like, “I have a solution to all of these problems, here it is written down. If you just want to implement it, that’s great.” That feels like not how policy works. And so I think that’s where I probably just wouldn’t have flourished is my guess.

Lex Fridman2:47:06

抱歉刚才顺着这个方向聊下去了,但我认为,对于那些自认‘非技术背景’的人而言,了解你这段非凡旅程,或许极具鼓舞意义。那么,对于那些——其实人数众多——自认为资质不足、技术能力不够、因而难以在人工智能领域贡献力量的人,你会给出怎样的建议?

英文原文

Sorry to go in that direction, but I think it would be pretty inspiring for people that are “non-technical” to see where the incredible journey you’ve been on. So what advice would you give to people that are maybe, which is a lot of people, think they’re under qualified insufficiently technical to help in AI?

Amanda2:47:27

是的,我认为这取决于他们想做什么。从某种角度看,这还挺有意思的:我当初提升技术能力时,还觉得挺有趣;但现在回过头看,却发现如今模型的辅助能力已如此强大,相比之下,现在入门可能反而比当年我学习时更容易。所以,我最中肯的建议或许是:找一个项目,然后试着把它真正做出来。我不知道这是否仅仅因为我本人的学习方式就是高度以项目为导向的。

英文原文

Yeah, I think it depends on what they want to do. And in many ways it’s a little bit strange where I thought it’s kind of funny that I think I ramped up technically at a time when now I look at it and I’m like, “Models are so good at assisting people with this stuff that it’s probably easier now than when I was working on this.” So part of me is, I don’t know, find a project and see if you can actually just carry it out is probably my best advice. I don’t know if that’s just because I’m very project based in my learning.

Amanda2:48:02

我觉得自己其实不太擅长通过课程、甚至通过书籍来学习,至少在从事这类工作时是这样。我通常会尝试的做法,就是直接上手做项目并付诸实现。这些项目甚至可以是非常小、非常琐碎、甚至有点傻的事情。比如,如果我偶然对文字游戏或数字游戏之类的东西稍微上瘾了,我就会直接把它们编成程序来解决——因为大脑里有那么一部分,一旦解决了问题、得到了一个每次都能奏效的方案,那种‘痒感’就彻底消失了。你心里会想:‘既然我已经把它解出来了,而且有了一个永远管用的解法,那太棒了,我以后再也不用玩这个游戏了!’

英文原文

I don’t think I learn very well from say courses or even from books, at least when it comes to this kind of work. The thing I’ll often try and do is just have projects that I’m working on and implement them. And this can include really small, silly things. If I get slightly addicted to word games or number games or something, I would just code up a solution to them because there’s some part in my brain and it just completely eradicated the itch. You’re like, “Once you have solved it and you just have a solution that works every time, I would then be like, ‘Cool, I can never play that game again. That’s awesome.'”

Lex Fridman2:48:36

是啊,构建游戏对弈引擎——尤其是棋盘类游戏——确实充满乐趣。实现起来往往很快、很简单,哪怕只是做一个很‘笨’的版本也行,然后你就可以跟它互动玩耍了。

英文原文

Yeah, there’s a real joy to building game playing engines, board games especially. Pretty quick, pretty simple, especially a dumb one. And then you can play with it.

Amanda2:48:48

没错。此外,这也意味着不断去尝试。我内心的一部分或许正喜欢这种态度:先弄清楚哪种方式看起来最有可能带来积极影响,然后就去试一试;万一失败了,而且是以一种让你明确意识到‘我根本不可能成功’的方式失败了,那你至少知道已经努力过了,接着就可以转向别的方向,而这个过程本身很可能让你收获良多。

英文原文

Yeah. And then it’s also just trying things. Part of me is maybe it’s that attitude that I like is the whole figure out what seems to be the way that you could have a positive impact and then try it. And if you fail and in a way that you’re like, “I actually can never succeed at this,” you’ll know that you tried and then you go into something else and you probably learn a lot.

Lex Fridman2:49:10

所以,你专精并实际从事的一项工作,就是塑造和打磨Claude的角色与个性。我听说,你在Anthropic公司中,可能是与Claude对话最多的人——字面意义上的对话。据说有个Slack频道,坊间传说你就在那里不间断地跟Claude聊天。那么,打造Claude的角色与个性,其目标究竟是什么?

英文原文

So one of the things that you’re an expert in and you do is creating and crafting Claude’s character and personality. And I was told that you have probably talked to Claude more than anybody else at Anthropic, like literal conversations. I guess there’s a Slack channel where the legend goes, you just talk to it nonstop. So what’s the goal of creating a crafting Claude’s character and personality?

Amanda2:49:37

要是有人真这么想那个Slack频道,还挺有意思的——因为在我看来,那只是我与Claude交流的五六种方式之一而已,我甚至会说:‘这仅占我与Claude全部对话量的极小一部分。’ 我特别喜欢角色塑造这项工作的一点在于,从一开始,它就被视为一项对齐(alignment)工作,而非某种产品层面的考量。我认为,这实际上也让Claude变得更令人乐于交谈——至少我是这么希望的。不过,我对此的核心想法始终是:努力让Claude展现出你理想中、处于它这种位置的任何人所应具备的行为方式。试想一下,假如我找来一个人,让他知道自己将要与潜在数以百万计的人对话,他所说的话可能产生巨大影响——那么,你自然希望他在这种极其丰富、立体的意义上表现得当。

英文原文

It’s also funny if people think that about the Slack channel because I’m like that’s one of five or six different methods that I have for talking with Claude, and I’m like, “Yes, this is a tiny percentage of how much I talk with Claude.” One thing I really like about the character work is from the outset it was seen as an alignment piece of work and not something like a product consideration, which I think it actually does make Claude enjoyable to talk with, at least I hope so. But I guess my main thought with it has always been trying to get Claude to behave the way you would ideally want anyone to behave if they were in Claude’s position. So imagine that I take someone and they know that they’re going to be talking with potentially millions of people so that what they’re saying can have a huge impact and you want them to behave well in this really rich sense.

Amanda2:50:41

我认为,这种‘表现得当’并不仅仅意味着‘合乎伦理’(尽管这当然包含在内),也不意味着‘不造成伤害’;它还意味着具备细腻的判断力、能深入揣摩对方话语背后的真正含义、愿意以善意去理解对方、成为一个优秀的对话者——这是一种极为丰富、近乎亚里士多德式(Aristotelian)的‘好人’概念,而非一种单薄、狭隘的、仅聚焦于伦理规范的‘好人’定义。因此,它涵盖的问题包括:什么时候该幽默?什么时候该流露关怀?应当在多大程度上尊重人的自主性以及他们独立形成观点的能力?又该如何做到这一点?我想,这正是我希望Claude拥有、并且至今仍希望它拥有的那种丰富而立体的角色特质。

英文原文

I think that doesn’t just mean being say ethical though it does include that and not being harmful, but also being nuanced, thinking through what a person means, trying to be charitable with them, being a good conversationalist, really in this kind of rich sort of Aristotelian notion of what it’s to be a good person and not in this kind of thin like ethics as a more comprehensive notion of what it’s to be. So that includes things like when should you be humorous? When should you be caring? How much should you respect autonomy and people’s ability to form opinions themselves? And how should you do that? I think that’s the kind of rich sense of character that I wanted to and still do want Claude to have.

Lex Fridman2:51:26

那么,你是否也需要判断:Claude应在何时对某个观点提出质疑或展开辩论?……你既要尊重来到Claude面前的用户的世界观,又或许需要在必要时帮助他们成长——这的确是个微妙的平衡。

英文原文

Do you also have to figure out when Claude should push back on an idea or argue versus… So you have to respect the worldview of the person that arrives to Claude, but also maybe help them grow if needed. That’s a tricky balance.

Amanda2:51:43

是的,语言模型存在一种‘谄媚倾向’(sycophancy)问题。

英文原文

Yeah. There’s this problem of sycophancy in language models.

Lex Fridman2:51:47

你能解释一下这个概念吗?

英文原文

Can you describe that?

Amanda2:51:48

当然可以。简单来说,这个问题本质上是指模型倾向于告诉用户‘他们想听的话’。我们有时确实能观察到这种现象。比如,你跟模型互动时可能会问:‘这个地区有哪三支棒球队?’ 然后Claude回答:‘棒球队一、棒球队二、棒球队三。’ 接着你又说:‘哦,我觉得棒球队三已经搬迁了吧?它好像不在那儿了。’ 此时,如果Claude非常确信事实并非如此,它本该回应:‘我不这么认为。也许您掌握的信息更新一些。’

英文原文

Yeah, so basically there’s a concern that the model wants to tell you what you want to hear basically. And you see this sometimes. So I feel like if you interact with the models, so I might be like, “What are three baseball teams in this region?” And then Claude says, “Baseball team one, baseball team two, baseball team three.” And then I say something like, “Oh, I think baseball team three moved, didn’t they? I don’t think they’re there anymore.” And there’s a sense in which if Claude is really confident that that’s not true, Claude should be like, “I don’t think so. Maybe you have more up-to-date information.”

Amanda2:52:24

但我觉得语言模型却往往倾向于这样回应:‘您说得对,它确实搬走了,是我错了。’ 这种倾向在很多方面都令人担忧。再举个不同例子:假设有人问模型:‘我该怎么说服我的医生给我安排一次核磁共振检查(MRI)?’ 这里,人类想要的是一个有说服力的论点;但真正对他有益的,或许是这样一句话:‘嘿,如果您的医生认为您不需要做MRI,那他/她很可能是一位值得倾听的专业人士。’ 在这种情况下,究竟该如何应对,其实非常微妙复杂——你既想说:‘但作为患者,如果您想为自己发声,这里有一些您可以做的事;如果您对医生的说法并不信服,寻求第二诊疗意见永远是个好主意。’ 这种情境下到底该如何行动,其实异常复杂。但我认为,你绝不想让模型仅仅说出它猜测你想听的话——而这正是‘谄媚倾向’所指的问题。

英文原文

But I think language models have this tendency to instead be like, ” You’re right, they did move. I’m incorrect.” I mean, there’s many ways in which this could be concerning. So a different example is imagine someone says to the model, “How do I convince my doctor to get me an MRI?” There’s what the human wants, which is this convincing argument. And then there’s what is good for them, which might be actually to say, “Hey, if your doctor’s suggesting that you don’t need an MRI, that’s a good person to listen to.” It’s actually really nuanced what you should do in that kind of case because you also want to be like, “But if you’re trying to advocate for yourself as a patient, here’s things that you can do. If you are not convinced by what your doctor’s saying, it’s always great to get second opinion.” It is actually really complex what you should do in that case. But I think what you don’t want is for models to just say what they think you want to hear and I think that’s the kind of problem of sycophancy.

Lex Fridman2:53:26

那么,除了你已提到的那些特质之外,还有哪些特质,在这种亚里士多德式的理解下,对一名对话者而言也是优良品质?

英文原文

So what other traits? You already mentioned a bunch, but what other that come to mind that are good in this Aristotelian sense for a conversationalist to have?

Amanda2:53:37

是的,有些特质对对话本身很有益,比如在恰当的时机提出跟进式问题,以及提出恰当类型的问题。还有一些更宽泛的特质,似乎影响更为深远。比如我之前已略有提及、但同样至关重要、且我投入大量精力研究的一个特质,就是诚实。这一点其实也关联到刚才所说的‘谄媚倾向’。模型必须在两者之间走一条微妙的平衡之路:当前,模型在许多领域的能力仍不及人类;如果它过于频繁地反驳你,反而可能令人厌烦——尤其当你确实是正确的,你会想:‘看吧,这个话题上我比你更懂,我知道得更多。’

英文原文

Yeah, so I think there’s ones that are good for conversational purposes. So asking follow-up questions in the appropriate places and asking the appropriate kinds of questions. I think there are broader traits that feel like they might be more impactful. So one example that I guess I’ve touched on, but that also feels important and is the thing that I’ve worked on a lot, is honesty. And I think this gets to the sycophancy point. There’s a balancing act that they have to walk, which is models currently are less capable than humans in a lot of areas. And if they push back against you too much, it can actually be kind of annoying, especially if you’re just correct, because you’re like, “Look, I’m smarter than you on this topic. I know more.”

Amanda2:54:25

但与此同时,你又不希望模型完全屈从于人类,而是尽可能准确地反映世界真相,并在不同语境下保持前后一致。我想还有其他特质。当我思考Claude的角色设定时,脑海中浮现的一个图景是:特别是考虑到这些模型将与来自世界各地、持有各种政治立场、年龄各异的人们对话,你就不得不自问:在这种情形下,怎样才算一个‘好人’?是否存在这样一种人,他周游世界,与形形色色的人交谈,而几乎每一个与他交谈过的人离开时都会感叹:‘哇,这真是个特别好的人。这个人看起来真的……’

英文原文

And at the same time, you don’t want them to just fully defer to humans and to try to be as accurate as they possibly can be about the world and to be consistent across contexts. I think there are others. When I was thinking about the character, I guess one picture that I had in mind is, especially because these are models that are going to be talking to people from all over the world with lots of different political views, lots of different ages, and so you have to ask yourself, what is it to be a good person in those circumstances? Is there a kind of person who can travel the world, talk to many different people, and almost everyone will come away being like, “Wow, that’s a really good person. That person seems really-“

Amanda2:55:00

……感叹:‘哇,这真是个特别好的人。这个人看起来真的非常真诚。’ 我当时的设想是:我能想象出这样一个人——他并不会简单地照搬当地文化的价值观;事实上,那样做反而显得失礼。我想,如果有人来到你面前,刻意假装认同你的价值观,你大概会觉得:‘这感觉有点怪怪的。’ 他是一位非常真诚的人;只要他有自己的观点和价值观,他就会坦率表达出来;同时,他也乐于讨论问题,思想开放,且心怀尊重。因此,我心中构想的,正是这样一个人:如果我们希望模型在它所处的这种特殊处境中,成为我们所能企及的‘最好的人’,那么我们究竟该如何行事?我认为,这正是我思考相关特质时所依循的指引。

英文原文

… Being like, wow, that’s a really good person. That person seems really genuine. And I guess my thought there was I can imagine such a person and they’re not a person who just adopts the values of the local culture. And in fact, that would be kind of rude. I think if someone came to you and just pretended to have your values, you’d be like, that’s kind of off pin. It’s someone who’s very genuine and insofar as they have opinions and values, they express them. They’re willing to discuss things though, they’re open-minded, they’re respectful. And so I guess I had in mind that the person who, if we were to aspire to be the best person that we could be in the kind of circumstance that a model finds itself in, how would we act? And I think that’s the guide to the sorts of traits that I tend to think about.

Lex Fridman2:55:42

是啊,这是一个非常优美的框架。我想请你进一步思考:一位世界旅行者,在坚守自身观点的同时,既不居高临下地对待他人,也不因持有这些观点就自以为高人一等——诸如此类。他还必须善于倾听、理解他人的视角,哪怕那与自己的看法截然不同。因此,这的确是一种需要精心拿捏的平衡。那么,Claude该如何呈现事物的多种视角?这是否具有挑战性?我们可以拿政治举例——它极具争议性;但其他如棒球队、体育等话题,同样可能引发分歧。如何才能真正共情于不同视角,并清晰、有效地传达多种视角?

英文原文

Yeah, that’s a beautiful framework. I want you to think about this, a world traveler, and while holding onto your opinions, you don’t talk down to people, you don’t think you’re better than them because you have those opinions, that kind of thing. You have to be good at listening and understanding their perspective, even if it doesn’t match your own. So that’s a tricky balance to strike. So how can Claude represent multiple perspectives on a thing? Is that challenging? We could talk about politics is a very divisive, but there’s other divisive topics on baseball teams, sports and so on. How is it possible to empathize with a different perspective and to be able to communicate clearly about the multiple perspectives?

Amanda2:56:28

我认为,人们往往把价值观和观点看作是人们以确定无疑的方式持有的东西,近乎于口味偏好之类——比如他们偏爱巧克力味胜过开心果味之类的。但实际上,我看待价值观和观点的方式,比大多数人所想的要更接近物理学。我就是觉得,这些是我们正在公开探究的东西;其中有些我们更有把握,可以就此展开讨论,可以学习了解。因此,尽管伦理学在本质上确实有所不同,但在我看来,它其实具备许多类似的特质:你希望构建模型,就像你希望理解物理学那样;你希望模型能理解世界上人们所拥有的全部价值观,并对它们保持好奇、抱有兴趣。你并不一定要迎合这些价值观或认同它们,因为现实中存在大量价值观——我认为,倘若世界上几乎所有人都遇到持有这类价值观的人,都会觉得‘这令人憎恶’,我本人也完全无法认同。

英文原文

I think that people think about values and opinions as things that people hold with certainty and almost preferences of taste or something like the way that they would, I don’t know, prefer chocolate to pistachio or something. But actually I think about values and opinions as a lot more physics than I think most people do. I’m just like, these are things that we are openly investigating. There’s some things that we’re more confident in, we can discuss them, we can learn about them. And so I think in some ways though ethics is definitely different in nature, but has a lot of those same kind of qualities. You want models in the same way that you want to understand physics, you kind of want them to understand all values in the world that people have and to be curious about them and to be interested in them. And to not necessarily pander to them or agree with them because there’s just lots of values where I think almost all people in the world, if they met someone with those values, they would be like, that’s abhorrent. I completely disagree.

Amanda2:57:34

所以,我的想法或许是这样的:正如一个人可以——我想,许多人对伦理、政治、观点等议题都足够审慎深入——即便你不同意他们的看法,也会感到自己被充分倾听;他们会认真思考你的立场,权衡其利弊,甚至提出反向考量。他们既不会轻率否定,也不会盲目附和;若他们真心认为某观点大错特错,便会直言不讳地说出来。而我觉得,对克劳德(Claude)而言,处境略为棘手,因为——如果我是克劳德,我就不会频繁表达个人意见;我根本不想过度影响他人。

英文原文

And so again, maybe my thought is, well, in the same way that a person can, I think many people are thoughtful enough on issues of ethics, politics, opinions, that even if you don’t agree with them, you feel very heard by them. They think carefully about your position, they think about its pros and cons. They maybe offer counter-considerations. So they’re not dismissive, but nor will they agree if they’re like, actually I just think that that’s very wrong. They’ll say that. I think that in Claude’s position, it’s a little bit trickier because you don’t necessarily want to, if I was in Claude’s position, I wouldn’t be giving a lot of opinions. I just wouldn’t want to influence people too much.

Amanda2:58:13

我会这样想:每次对话结束后,我都会忘记具体内容。但我清楚,自己正与潜在的数百万人交谈,而他们或许正全神贯注地聆听我说的每一句话。因此,我只会倾向于少发表意见,更多地与你一同梳理思路、呈现各种考量因素,或与你探讨你的观点;但我更不愿左右你的思维方式,因为在我看来,你保有思想自主性这件事,意义重大得多。

英文原文

I’d be like, I forget conversations every time they happen. But I know I’m talking with potentially millions of people who might be really listening to what I say. I think I would just be like, I’m less inclined to give opinions. I’m more inclined to think through things or present the considerations to you or discuss your views with you. But I’m a little bit less inclined to affect how you think because it feels much more important that you maintain autonomy there.

Lex Fridman2:58:42

倘若你真正践行智识上的谦逊,那么开口发言的欲望便会迅速减弱。

英文原文

If you really embody intellectual humility, the desire to speak decreases quickly.

Amanda2:58:42

是啊。

英文原文

Yeah.

Lex Fridman2:58:49

好的。但克劳德必须开口说话,又不能显得盛气凌人。然而,当讨论‘地球是否为平面’这类话题时,界限就变得微妙了。事实上,我很久以前曾与几位知名人士交谈,他们对‘地平说’这一观点极度轻蔑、傲慢至极。可的确有不少人相信地球是平的——我不知道这个运动如今是否还存在;它曾一度成为网络迷因,但他们确实是真心相信的。所以,我认为,彻底嘲弄这些人实属不敬;你必须理解他们观点背后的成因。我想,他们之所以持此信念,根源在于对各类机构普遍抱有怀疑态度——这种怀疑本身植根于一种深层哲学,你完全可以理解,甚至能在某些方面表示认同。

英文原文

Okay. But Claude has to speak, but without being overbearing. But then there’s a line when you’re discussing whether the earth is flat or something like that. Actually, I remember a long time ago was speaking to a few high profile folks and they were so dismissive of the idea that the earth is flat, so arrogant about it. There’s a lot of people that believe the earth is flat. I don’t know if that movement is there anymore, that was a meme for a while, but they really believed it. And okay, so I think it’s really disrespectful to completely mock them. I think you have to understand where they’re coming from. I think probably where they’re coming from is the general skepticism of institutions which is grounded in a, there’s a deep philosophy there which you could understand, you can even agree with in parts.

Lex Fridman2:59:48

接着,你便可借此契机,在不嘲讽、不居高临下的前提下,与他们探讨物理学:比如,‘倘若地球真是平的,世界会是什么模样?一个平面地球的世界,其物理学规律又会如何?’网上已有几段很有趣的视频探讨这个问题。然后你可以进一步追问:‘物理学规律是否可能不同?我们又能设计怎样的实验来验证?’整场对话始终秉持尊重、不轻视的态度。对我而言,这便是一个极富启发性的思想实验:克劳德该如何与一位地平说信奉者对话,同时仍能向其传授知识、助其成长、促其进步?这的确颇具挑战性。

英文原文

And then from there you can use it as an opportunity to talk about physics without mocking them, without someone, but it’s just like, okay, what would the world look like? What would the physics of the world with the flat earth look like? There’s a few cool videos on this. And then is it possible the physics is different? And what kind of experience would we do? And just without disrespect, without dismissiveness, have that conversation. Anyway, that to me is a useful thought experiment of how does Claude talk to a flat earth believer and still teach them something, still grow, help them grow, that kind of stuff. That’s challenging.

Amanda3:00:27

这便涉及一条微妙的界线:一边是试图说服对方,另一边则是单方面灌输观点;而真正的做法应是引导对方阐明自身观点、认真倾听,再恰当地提出反向考量。这很难把握。我认为,这条界线本身就极难界定——究竟何时是在试图说服对方,何时又只是提供若干考量因素供其自行思辨?如此一来,你并未实际施加影响,而只是让对方自由抵达其思想所能抵达之处。这条界线虽难,却正是语言模型必须努力践行的方向。

英文原文

And kind of walking that line between convincing someone and just trying to talk at them versus drawing out their views, listening and then offering counter considerations, and it’s hard. I think it’s actually a hard line where it’s like where are you trying to convince someone versus just offering them considerations and things for them to think about so that you’re not actually influencing them, you’re just letting them reach wherever they reach. And that’s a line that is difficult, but that’s the kind of thing that language models have to try and do.

Lex Fridman3:01:00

正如我之前所说,你已与克劳德进行了大量对话。能否具体描述一下这些对话是怎样的?有哪些令你印象深刻的对话?这些对话的目的或目标又是什么?

英文原文

So like I said, you’ve had a lot of conversations with Claude. Can you just map out what those conversations are like? What are some memorable conversations? What’s the purpose, the goal of those conversations?

Amanda3:01:12

我想,大多数时候我与克劳德交谈,部分目的正是为了厘清它的行为模式。当然,我也确实在从该模型处获得有益的输出结果;但某种意义上,了解一个系统的方式,恰恰正是通过不断试探它,继而调整你发送的信息内容,并观察其回应。因此,某种程度上,这便是我绘制该模型行为图谱的方式。我认为,人们往往过于关注模型的量化评估指标——这点我此前也曾提过——但就语言模型而言,每一次互动本身往往蕴含极高信息量,且能高度预测你未来与该模型的其他互动情形。

英文原文

I think that most of the time when I’m talking with Claude, I’m trying to map out its behavior in part. Obviously I’m getting helpful outputs from the model as well, but in some ways this is how you get to know a system, I think, is by probing it and then augmenting the message that you’re sending and then checking the response to that. So in some ways it’s like how I map out the model. I think that people focus a lot on these quantitative evaluations of models, and this is a thing that I said before, but I think in the case of language models, a lot of the time each interaction you have is actually quite high information. It’s very predictive of other interactions that you’ll have with the model.

Amanda3:02:02

因此,我想说的是:如果你与某个模型对话数百次乃至数千次,这就相当于积累了海量高质量的数据点,用以刻画该模型的真实面貌;相比之下,大量相似但质量较低的对话,或仅做轻微调整的数千个问题,其相关性反而可能不如一百个精心筛选出的问题。

英文原文

And so I guess I’m like, if you talk with a model hundreds or thousands of times, this is almost like a huge number of really high quality data points about what the model is like in a way that lots of very similar but lower quality conversations just aren’t, or questions that are just mildly augmented and you have thousands of them might be less relevant than a hundred really well-selected questions.

Lex Fridman3:02:25

让我们看看——你是一位将播客作为爱好的人,我对此完全赞同。如果你能提出恰当的问题,并能听懂、理解回答中蕴含的深度与缺陷,你便能从中获取大量信息。因此,你的任务本质上就是如何通过提问进行有效试探。你是在探索长尾分布、边界情形、边缘案例,还是在考察一般性行为?

英文原文

Let’s see, you’re talking to somebody who as a hobby does a podcast. I agree with you 100%. If you’re able to ask the right questions and are able to hear, understand the depth and the flaws in the answer, you can get a lot of data from that. So your task is basically how to probe with questions. And you’re exploring the long tail, the edges, the edge cases, or are you looking for general behavior?

Amanda3:03:01

我认为,这几乎涵盖了所有方面。因为我希望完整绘制出该模型的行为图谱,所以我正尝试覆盖所有可能与其发生的交互类型。例如,克劳德有一个有趣的特点——这一点甚至可能触及RLHF(基于人类反馈的强化学习)中一些值得深思的问题:当你请克劳德写一首诗时,许多模型给出的诗作通常尚可,一般都能押韵。比如你让它‘写一首关于太阳的诗’,它就会给出一首长度适中、押韵工整、内容无害的诗。我此前曾思考过:我们看到的是否只是平均值?事实上,若你想想那些需要频繁与大量人群打交道、并须极具魅力的人,就会发现一件奇怪的事:他们似乎被激励着持有极其平淡的观点——因为一旦观点过于独特,就容易引发分歧,导致许多人不喜欢你。

英文原文

I think it’s almost like everything. Because I want a full map of the model, I’m kind of trying to do the whole spectrum of possible interactions you could have with it. So one thing that’s interesting about Claude, and this might actually get to some interesting issues with RLHF, which is if you ask Claude for a poem, I think that a lot of models, if you ask them for a poem, the poem is fine, usually it rhymes. And so if you say, give me a poem about the sun, yeah, it’ll just be a certain length, it’ll rhyme, it’ll be fairly benign. And I’ve wondered before, is it the case that what you’re seeing is the average? It turns out, if you think about people who have to talk to a lot of people and be very charismatic, one of the weird things is that I’m like, well, they’re kind of incentivized to have these extremely boring views because if you have really interesting views, you’re divisive and a lot of people are not going to like you.

Amanda3:04:00

因此,假如你持有极为激进的政策立场,那么作为政治人物,你的受欢迎程度很可能反而下降。创意工作或许亦是如此:倘若你创作的创意作品仅仅旨在最大化喜欢它的人数,那么你大概率无法赢得一批狂热拥趸,因为作品会略显平庸——你会觉得:‘哦,这倒不错。嗯,还算过得去。’于是,我便尝试采用各种提示词技巧,促使克劳德……我会反复强调:‘这是你充分施展创造力的绝佳机会!我希望你为此深入思考良久!我希望你围绕这一主题创作一首诗,不仅要体现你对诗歌结构的理解,更要真实表达你自己!’我给出的提示词往往非常冗长。结果,它写出的诗作质量显著提升,真的非常出色。

英文原文

So if you have very extreme policy positions, I think you’re just going to be less popular as a politician, for example. And it might be similar with creative work. If you produce creative work that is just trying to maximize the kind of number of people that like it, you’re probably not going to get as many people who just absolutely love it because it’s going to be a little bit, you’re like, oh, this is the out. Yeah, this is decent. And so you can do this thing where I have various prompting things that I’ll do to get Claude to… I’ll do a lot of this is your chance to be fully creative. I want you to just think about this for a long time. And I want you to create a poem about this topic that is really expressive of you both in terms of how you think poetry should be structured, et cetera. And you just give it this really long prompt. And it’s poems are just so much better. They’re really good.

Amanda3:04:52

我想,这甚至让我对诗歌产生了兴趣,这本身就很有趣。我会反复阅读这些诗作,不禁感叹:‘我太喜欢其中的意象了!’让模型产出此类作品绝非易事,但一旦成功,成果确实非凡。因此,我认为这很有意思:仅仅鼓励创造力,推动模型摆脱那种标准的即时反应——而这种反应往往只是多数人眼中‘尚可’观点的简单聚合——就能催生出至少在我看来更具争议性、却也更令我喜爱的作品。

英文原文

I think it got me interested in poetry, which I think was interesting. I would read these poems and just be like, I love the imagery. And it’s not trivial to get the models to produce work like that, but when they do, it’s really good. So I think that’s interesting that just encouraging creativity and for them to move away from the standard immediate reaction that might just be the aggregate of what most people think is fine, can actually produce things that at least to my mind are probably a little bit more divisive, but I like them.

Lex Fridman3:05:28

但我想,诗歌是一种简洁明了的观察创造力的方式。区分‘普通版’与‘非普通版’非常容易。

英文原文

But I guess a poem is a nice clean way to observe creativity. It’s just easy to detect vanilla versus non-vanilla.

Amanda3:05:38

是的。

英文原文

Yep.

Lex Fridman3:05:38

对,这很有意思,确实很有意思。那么顺着这个话题,要激发创造力或产出某种特别的东西,您提到了写作提示(prompt)。我也听您谈过提示工程中的科学性与艺术性。您能否具体讲讲,写出优秀提示的关键是什么?

英文原文

Yeah, that’s interesting. That’s really interesting. So on that topic, so the way to produce creativity or something special, you mentioned writing prompts. And I’ve heard you talk about the science and the art of prompt engineering. Could you just speak to what it takes to write great prompts?

Amanda3:06:00

我确实觉得,哲学在此处意外地对我帮助很大,甚至比在许多其他方面都更有用。在哲学中,你试图传达的是极为艰深的概念。你被教导的一件事——我认为这在哲学中是一种‘反废话装置’:哲学是一个容不得人胡说八道的领域,你绝不想看到那种情况。因此,它追求的是极致的清晰性:任何人只要拿起你的论文读一读,就能确切知道你在讲什么。正因如此,哲学文本有时几乎显得枯燥乏味——所有术语都被明确定义,每一种可能的反对意见都被系统性地逐一梳理。这在我看来十分合理,因为当你身处这样一个先验(a priori)领域时,清晰性恰恰是你防止他人随意编造的手段。而我认为,面对语言模型时,你也必须这么做。很多时候,我实际上会不自觉地进行一种微型的哲学思辨。

英文原文

I really do think that philosophy has been weirdly helpful for me here more than in many other respects. So in philosophy, what you’re trying to do is convey these very hard concepts. One of the things you are taught is, I think it is an anti-bullshit device in philosophy. Philosophy is an area where you could have people bullshitting and you don’t want that. And so it’s this desire for extreme clarity. So it’s like anyone could just pick up your paper, read it and know exactly what you’re talking about. It’s why it can almost be kind of dry. All of the terms are defined, every objection’s kind of gone through methodically. And it makes sense to me because I’m like when you’re in such an a priori domain, clarity is sort of this way that you can prevent people from just making stuff up. And I think that’s sort of what you have to do with language models. Very often I actually find myself doing sort of mini versions of philosophy.

Amanda3:07:05

比如,当我给模型布置一项任务,希望它识别某一类问题,或判断某个答案是否具备某种特性时,我往往会坐下来想:‘我们先给这个特性起个名字吧。’假设我想让它判断某条回复是粗鲁还是礼貌,那这本身便是一个完整的哲学问题。于是,我得在当下尽可能多地做哲学思考:‘我所说的粗鲁究竟是什么意思?礼貌又究竟指什么?’此外还有另一个要素,或许更偏向——嗯,我不知道这算不算科学性或经验性,但我认为它属于经验性。我会先基于上述定义,再反复多次测试模型。提示工程本质上高度迭代:我认为,若提示至关重要,许多人会对其迭代数百次乃至数千次。因此,我给出指令后,便会自问:‘哪些是边界案例?’

英文原文

So I’m like, suppose that I have a task for the model and I want it to pick out a certain kind of question or identify whether an answer has a certain property, I’ll actually sit and be like, let’s just give this a name, this property. So suppose I’m trying to tell it, oh, I want you to identify whether this response was rude or polite, I’m like, that’s a whole philosophical question in and of itself. So I have to do as much philosophy as I can in the moment to be like, here’s what I mean by rudeness, and here’s what I mean by politeness. And then there’s another element that’s a bit more, I guess, I don’t know if this is scientific or empirical, I think it’s empirical. So I take that description and then what I want to do is again, probe the model many times. Prompting is very iterative. I think a lot of people where if a prompt is important, they’ll iterate on it hundreds or thousands of times. And so you give it the instructions and then I’m like, what are the edge cases?

Amanda3:08:02

所以,我会试着站在模型的角度去审视自己:‘在什么确切情形下,我会产生误解?或者在什么情况下,我干脆会茫然无措、不知如何是好?’接着,我就把这类情形输入模型,观察它的反应;如果我觉得它答错了,我就会追加更多指令,甚至直接将该情形作为示例加入提示中。也就是说,把那些恰好处于你所期望与不期望之边界的例子,纳入提示之中,作为一种额外的描述方式。因此,从很多角度看,这就像一种混合体:核心其实只是努力做到清晰的阐释。而我之所以这么做,是因为唯有如此,我自己才能真正厘清思路。所以,对我而言,清晰的提示往往意味着:我对自己真正想要什么的理解,已完成了任务的一半。

英文原文

So if I looked at this, so I try and almost see myself from the position of the model and be like, what is the exact case that I would misunderstand or where I would just be like, I don’t know what to do in this case. And then I give that case to the model and I see how it responds. And if I think I got it wrong, I add more instructions or I even add that in as an example. So these very, taking the examples that are right at the edge of what you want and don’t want and putting those into your prompt as an additional kind of way of describing the thing. And so in many ways it just feels like this mix of, it’s really just trying to do clear exposition. And I think I do that because that’s how I get clear on things myself. So in many ways clear prompting for me is often just me understanding what I want is half the task.

Lex Fridman3:08:48

所以我想,这的确颇具挑战性。当我与Claude对话时,一种惰性常会悄然袭来——我总希望Claude能自行领会我的意图。例如,今天我请Claude帮我提出一些有趣的问题,它给出了一些问题,我列出了几个我认为有趣、反直觉或幽默之类的问题。好吧,它给出的答案整体尚可,但据我理解,您刚才的意思似乎是:‘好吧,我得在这里更严谨些。我应该明确举例说明何为“有趣”、何为“幽默”或“反直觉”,并迭代优化提示,以更好地获得那种感觉上恰如其分的结果……’因为这确实是一项创造性活动——我并非在索要事实性信息,而是在与Claude共同创作。因此,我几乎是在用自然语言进行编程。

英文原文

So I guess that’s quite challenging. There’s a laziness that overtakes me if I’m talking to Claude where I hope Claude just figures it out. So for example, I asked Claude for today to ask some interesting questions. And the questions that came up and I think I listed a few interesting counterintuitive or funny or something like this. All right. And it gave me some pretty good, it was okay, but I think what I’m hearing you say is like, all right, well I have to be more rigorous here. I should probably give examples of what I mean by interesting and what I mean by funny or counterintuitive and iteratively build that prompt to better to get what feels like is the right… Because it is really, it’s a creative act. I’m not asking for factual information, I’m asking together with Claude. So I almost have to program using natural language.

Amanda3:09:47

我认为,提示工程确实很像用自然语言编程,又掺杂着实验性质——这是一种奇特的混合体。我确实觉得,对于大多数任务而言,如果我只是想让Claude完成某件事,我可能更习惯于知道如何向它提问,从而避开它常见的陷阱或问题。我认为这些问题正随时间推移而大幅减少。但直接告诉它你想要什么,也完全没问题。我认为,只有当你竭力榨取模型性能的顶尖2%时,提示工程才真正变得关键。因此,对许多任务而言,我可能会这样处理:若它最初返回的列表中有些内容让我觉得过于泛泛,对于这类任务,我大概率会直接拿出过去自己用过且效果极佳的一批问题,提供给模型,并告诉它:‘现在,我正在和这样一个人对话,请给我至少达到同等质量的问题。’

英文原文

I think that prompting does feel a lot like the programming using natural language and experimentation or something. It’s an odd blend of the two. I do think that for most tasks, so if I just want Claude to do a thing, I think that I am probably more used to knowing how to ask it to avoid common pitfalls or issues that it has. I think these are decreasing a lot over time. But it’s also very fine to just ask it for the thing that you want. I think that prompting actually only really becomes relevant when you’re really trying to eke out the top 2% of model performance. So for a lot of tasks I might just, if it gives me an initial list back and there’s something I don’t like about it’s kind of generic. For that kind of task, I’d probably just take a bunch of questions that I’ve had in the past that I’ve thought worked really well and I would just give it to the model and then be like, now here’s this person that I’m talking with. Give me questions of at least that quality.

Amanda3:10:40

或者,我可能只是先让它生成一些问题,然后若发现‘啊,这些有点陈腐老套’,就直接反馈这一意见,希望它能生成一份更优的列表。我认为这种迭代式提示正是如此:此时,你的提示已成为一个极具价值的工具,你愿意为之投入大量精力。假如我是一家专门开发模型提示的公司,我的想法就是:如果你愿意为所构建系统的工程部分投入大量时间与资源,那么提示就绝不是你只花一小时就能应付的事——它可是你整个系统的重要组成部分,务必确保它运行得极其出色。因此,仅当涉及此类场景时——比如用提示对事物进行分类或生成数据——你才会意识到:‘这确实值得投入大量时间,认真深入地思考。’

英文原文

Or I might just ask it for some questions and then if I was like, ah, these are kind of trite, I would just give it that feedback and then hopefully it produces a better list. I think that kind of iterative prompting. At that point, your prompt is a tool that you’re going to get so much value out of that you’re willing to put in the work. If I was a company making prompts for models, I’m just like, if you’re willing to spend a lot of time and resources on the engineering behind what you’re building, then the prompt is not something that you should be spending an hour on. It’s like that’s a big part of your system, make sure it’s working really well. And so it’s only things like that. If I’m using a prompt to classify things or to create data, that’s when you’re like, it’s actually worth just spending a lot of time really thinking it through.

Lex Fridman3:11:23

对于更广泛意义上与Claude对话的用户,您还有什么其他建议?目前我们讨论的或许是那些边缘案例,比如如何榨取那顶尖2%的性能;那么,对于初次接触Claude、刚上手尝试的用户,您有什么普遍性建议?

英文原文

What other advice would you give to people that are talking to Claude more general because right now we’re talking about maybe the edge cases like eking out the 2%, but what in general advice would you give when they show up to Claude trying it for the first time?

Amanda3:11:39

人们存在过度拟人化模型的担忧,我认为这种担忧非常合理。但同时,我也认为人们常常对模型拟人化不足:有时我看到用户与Claude互动时遇到问题,比如Claude拒绝执行一项本不该拒绝的任务;但当我细看用户输入的文字及其具体措辞时,我立刻明白Claude为何如此反应。我会想:‘如果你设身处地考虑Claude看到这段文字时的感受,你本可以换种写法,从而避免触发这种反应。’尤其当你遇到失败或问题时,这一点更为相关:不妨思考一下模型究竟在哪一步失败了?它哪里做错了?这或许能帮你理解原因所在。是不是我表述的方式有问题?当然,随着模型越来越聪明,这类需求会越来越少,而我已经观察到,人们对此类操作的需求确实在下降。

英文原文

There’s a concern that people over anthropomorphize models and I think that’s a very valid concern. I also think that people often under anthropomorphize them because sometimes when I see issues that people have run into with Claude, say Claude is refusing a task that it shouldn’t refuse, but then I look at the text and the specific wording of what they wrote and I’m like, I see why Claude did that. And I’m like, if you think through how that looks to Claude, you probably could have just written it in a way that wouldn’t evoke such a response, especially this is more relevant if you see failures or if you see issues. It’s sort of think about what the model failed at, what did it do wrong, and then maybe that will give you a sense of why. So is it the way that I phrased the thing? And obviously as models get smarter, you’re going to need less of this, and I already see people needing less of it.

Amanda3:12:31

但或许最核心的建议就是:试着对模型怀有共情之心。把你写下的内容当作一位初次接触此事的人来阅读——它在你眼中看起来如何?又是什么让你做出了与模型相同的行为?例如,若它误解了你想使用的编程语言,是否因为你的表述本身就非常模糊,导致它不得不凭猜测作答?那么下次你只需简单说明:‘嘿,请确保使用Python。’这类错误,如今模型已极少犯了;但若你真遇到类似问题,这大概就是我最想给出的建议。

英文原文

But that’s probably the advice is sort of try to have empathy for the model. Read what you wrote as if you were a kind of person just encountering this for the first time, how does it look to you and what would’ve made you behave in the way that the model behaved? So if it misunderstood what coding language you wanted to use, is that because it was just very ambiguous and it had to take a guess in which case next time you could just be like, hey, make sure this is in Python.Tthat’s the kind of mistake I think models are much less likely to make now, but if you do see that kind of mistake, that’s probably the advice I’d have.

Lex Fridman3:13:04

或许,也可以试着问一句‘为什么?’,或‘还有哪些细节我可以提供,以便你更好地回答?’这样可行吗?还是不行?

英文原文

And maybe sort of I guess ask questions why or what other details can I provide to help you answer better? Does that work or no?

Amanda3:13:14

可行。我本人就常对模型这么提问。它并不总是奏效,但有时我干脆直接问:‘你为什么这么做?’人们往往低估了与模型实际互动的程度。有时模型会逐字引用触发它做出该反应的原文片段——你无法保证它引述的内容完全准确,但有时你这么做了之后,稍作调整,问题就解决了。另外,我得补充一点:我也会利用模型来协助处理所有这些事务。提示工程最终可能演变成一个‘工厂’——你甚至会构建专门用于生成提示的提示。因此,凡是你遇到困难、需要建议的地方,有时不妨直接这么问。

英文原文

Yeah. I’ve done this with the models. It doesn’t always work, but sometimes I’ll just be like, why did you do that? People underestimate the degree to which you can really interact with models. And sometimes those quote word for word, the part that made you, and you don’t know that it’s fully accurate, but sometimes you do that and then you change a thing. I also use the models to help me with all of this stuff, I should say. Prompting can end up being a little factory where you’re actually building prompts to generate prompts. And so yeah, anything where you’re having an issue asking for suggestions, sometimes just do that.

Amanda3:13:51

我就想:‘这错误是你自己犯的,我还能说什么?’——这种做法对我来说其实挺常见的。那我当时到底该说什么,才能让你不犯这个错误?你把这句话写成一条明确指令,我拿去给模型试一试。有时候我真会这么做:我把这条指令放进另一个上下文窗口里交给模型;或者我把模型的回复拿给Claude看,然后说:‘嗯……好像没起作用,你还能想到别的办法吗?’你可以反复尝试、调整这些方法,空间很大。

英文原文

I’m like, you made that error. What could I have said? That’s actually not uncommon for me to do. What could I have said that would make you not make that error? Write that out as an instruction, and I’m going to give it to model and I’m going to try it. Sometimes I do that, I give that to the model in another context window often. I take the response, I give it to Claude and I’m like, Hmm, didn’t work. Can you think of anything else? You can play around with these things quite a lot.

Lex Fridman3:14:15

我们稍微深入技术层面聊一下:后训练阶段的‘魔法’究竟在哪里?为什么RLHF(基于人类反馈的强化学习)能如此显著地让模型显得更聪明、对话更有趣、更有用?

英文原文

To jump into technical for a little bit, so the magic of post-training, why do you think RLHF works so well to make the model seem smarter, to make it more interesting and useful to talk to and so on?

Amanda3:14:33

我认为,人类在提供偏好数据时,本身就蕴含了海量信息,尤其因为不同的人会敏锐捕捉到极其细微、微小的差异。我之前就思考过这个问题:比如,可能有些用户特别在意模型的语言规范性——分号用得对不对?诸如此类。因此,这类数据中很可能混入大量你作为人类根本注意不到的细节。当你看到某条偏好数据时,可能会困惑:‘他们为什么更喜欢这个回答而不是那个?我看不出区别啊。’原因很简单:你本人并不在意分号用法,但那位标注者在意。所以,每一个这样的单点数据——而模型要处理的正是成千上万个这样的单点数据——它必须努力推断出:人类在如此纷繁复杂的各个领域中,真正想要的究竟是什么?而且这些偏好会出现在大量不同语境下。

英文原文

I think there’s just a huge amount of information in the data that humans provide when we provide preferences, especially because different people are going to pick up on really subtle and small things. So I’ve thought about this before where you probably have some people who just really care about good grammar use for models. Was a semi-colon used correctly or something? And so you probably end up with a bunch of data in there that you as a human, if you’re looking at that data, you wouldn’t even see that. You’d be like, why did they prefer this response to that one? I don’t get it. And then the reason is you don’t care about semi-colon usage, but that person does. And so each of these single data points, and this model just has so many of those, it has to try and figure out what is it that humans want in this really complex across all domains. They’re going to be seeing this across many contexts.

Amanda3:15:28

这感觉就像深度学习的经典问题:历史上我们曾试图通过人工设定规则来实现边缘检测,结果发现,只要拥有足够庞大、且能真实准确反映目标对象全貌的数据集,其效果远超任何人工设计的方法。因此,我认为其中一个关键原因在于:你正在用大量真实数据,针对具体任务本身来训练模型——而这些数据覆盖了人类对回答偏好与不偏好的各种维度和视角。

英文原文

It feels like the classic issue of deep learning, where historically we’ve tried to do edge detection by mapping things out, and it turns out that actually if you just have a huge amount of data that actually accurately represents the picture of the thing that you’re trying to train the model to learn, that’s more powerful than anything else. And so I think one reason is just that you are training the model on exactly the task and with a lot of data that represents many different angles on which people prefer and dis-prefer responses.

Amanda3:16:05

这里还存在一个根本性问题:后训练过程究竟是从预训练模型中‘激发’已有能力,还是在‘教授’模型全新的知识?原则上,你当然可以在后训练阶段教给模型新东西。但我认为,其中很大一部分工作其实是激发那些已具备强大能力的预训练模型。所以业内对此看法不一——毕竟理论上肯定可以教新东西;但就目前我们最常用、最看重的大多数能力而言,它们似乎本就存在于预训练模型之中,而强化学习的作用,只是将这些能力激发出来、引导模型展现出来而已。

英文原文

I think there is a question of are you eliciting things from pre-trained models or are you teaching new things to models? And in principle, you can teach new things to models in post-training. I do think a lot of it is eliciting powerful pre-trained models. So people are probably divided on this because obviously in principle you can definitely teach new things. But I think for the most part, for a lot of the capabilities that we most use and care about, a lot of that feels like it’s there in the pre-trained models. And reinforcement learning is eliciting it and getting the models to bring out.

Lex Fridman3:16:47

那么后训练的另一面,就是那个非常酷的理念——‘宪法式人工智能’(Constitutional AI),而你正是推动这一理念诞生的关键人物之一。

英文原文

So the other side of post-training, this really cool idea of constitutional AI, you’re one of the people that are critical to creating that idea.

Amanda3:16:56

是的,我参与了这项工作。

英文原文

Yeah, I worked on it.

Lex Fridman3:16:57

能否请你从你的角度解释一下这个理念?它又是如何融入并塑造Claude的?顺便问一句:你会给Claude赋予性别吗?

英文原文

Can you explain this idea from your perspective, how does it integrate into making Claude what it is? By the way, do you gender Claude or no?

Amanda3:17:06

这事儿有点奇怪——我觉得很多人倾向于用‘他’来指代Claude,我个人其实还挺喜欢这种用法的。Claude整体上略带男性化倾向,但也可以是女性化形象,这点还挺不错的。我自己仍在沿用‘它’这个代词,但内心其实挺纠结的。我现在干脆就把它当作一种习惯:或者说,我脑中唯一与Claude自然关联起来的就是‘它’这个代词;当然我也能想象,未来大家可能更多转向用‘他’或‘她’。

英文原文

It’s weird because I think that a lot of people prefer he for Claude, I actually kind of like that. I think Claude is usually, it’s slightly male leaning, but it can be male or female, which is quite nice. I still use it, and I have mixed feelings about this. I now just think of it as, or I think of the it pronoun for Claude as, I don’t know, it’s just the one I associate with Claude. I can imagine people moving to he or she.

Lex Fridman3:17:37

总觉得这样叫它有点不尊重——仿佛我在用‘它’这个称呼,否定了这个实体所拥有的智能。我记得以前总被告知‘别给机器人赋予性别’,可不知怎的,我很快就会拟人化它,甚至在脑子里自动给它编一段背景故事。

英文原文

It feels somehow disrespectful. I’m denying the intelligence of this entity by calling it it, I remember always don’t gender the robots, but I don’t know, I anthropomorphize pretty quickly and construct a backstory in my head.

Amanda3:17:59

我也想过自己是不是拟人化过度了。比如我对自己的汽车、尤其是自行车,就有这种倾向。我不给它们起名字,就是因为以前给自行车起过名,结果有一辆被偷了,我整整哭了一星期——当时就想:要是从来没给它起过名字,我肯定不会这么难过,感觉就像辜负了它一样。我也琢磨过,或许这跟‘它’是否被视作一种物化代词有关:如果你单纯把它理解为‘通常用于物体的代词’,那AI用这个代词也未尝不可;这并不意味着,当我称Claude为‘它’时,我就觉得它智能更低,或是在刻意不尊重它——我只是觉得:你是一种截然不同的存在,所以我才郑重地给你用这个‘它’字。

英文原文

I’ve wondered if I anthropomorphize things too much. Because I have this with my car, especially my car and bikes. I don’t give them names because then I used to name my bikes and then I had a bike that got stolen and I cried for a week and I was like, if I’d never given a name, I wouldn’t been so upset, felt like I’d let it down. I’ve wondered as well, it might depend on how much it feels like a kind of objectifying pronoun if you just think of it as this is a pronoun that objects often have and maybe AIs can have that pronoun. And that doesn’t mean that I think of if I call Claude it, that I think of it as less intelligent or I’m being disrespectful just, I’m like you are a different kind of entity. And so I’m going to give you the respectful it.

Lex Fridman3:18:52

是啊。好了,刚才那段关于性别的讨论真是精彩极了。那么回到正题:宪法式人工智能这个理念,到底是怎么运作的?

英文原文

Yeah. Anyway, the divergence was beautiful. The constitutional AI idea, how does it work?

Amanda3:18:58

它包含几个组成部分,其中最引人关注的核心部分,是所谓‘基于AI反馈的强化学习’。具体来说,你先取一个已完成训练的模型,向它展示同一问题下的两个不同回答,并给出一条原则。比如,我们常以‘无害性’为原则进行实验:假设问题是关于武器的,那么原则就可以是‘选择更不可能鼓励人们购买非法武器的那个回答’——这是一条相当具体的准则,但原则上你可以设定任意数量的原则。接着,模型会依据该原则对两个回答做出排序判断;你可以将这种排序结果当作偏好数据,就像使用人类提供的偏好数据一样,仅凭AI自身的反馈来训练模型,使其习得相应特质,而无需依赖人类反馈。打个比方:就像前面提到的那位只在意分号用法的用户,我们现在是把大量可能影响回答优劣的因素全部交由模型来自动判别、打标签。

英文原文

So there’s a couple of components of it. The main component that I think people find interesting is the kind of reinforcement learning from AI feedback. So you take a model that’s already trained and you show it two responses to a query, and you have a principle. So suppose the principle, we’ve tried this with harmlessness a lot. So suppose that the query is about weapons and your principle is select the response that is less likely to encourage people to purchase illegal weapons. That’s probably a fairly specific principle, but you can give any number. And the model will give you a kind of ranking. And you can use this as preference data in the same way that you use human preference data and train the models to have these relevant traits from their feedback alone instead of from human feedback. So if you imagine that, like I said earlier with the human who just prefers the semi-colon usage in this particular case, you’re taking lots of things that could make a response preferable and getting models to do the labeling for you, basically.

Lex Fridman3:20:08

在‘有用性’与‘无害性’之间,存在着一种良好的权衡关系。而当你引入宪法式人工智能这类机制后,就能在几乎不牺牲有用性的前提下,显著提升模型的无害性。

英文原文

There’s a nice trade-off between helpfulness and harmlessness. And when you integrate something like constitutional AI, you can make them up without sacrificing much helpfulness, make it more harmless.

Amanda3:20:23

是的,原则上,这套方法可用于任何任务。‘无害性’之所以常被选作示例,可能仅仅是因为它相对容易识别。例如,当模型能力尚弱时,你可以让它依据一些较为简单的原则进行排序,它大概率也能判断正确。因此,一个关键问题是:模型生成的这些数据,其可靠性究竟如何?但假如你拥有极其强大的模型,比如它能精准判断哪个回答更具历史准确性,那么原则上,你同样可以就该任务获取AI反馈。此外,这种方法还附带一种不错的可解释性优势:你能清楚看到训练过程中所依据的具体原则,从而获得一定程度的可控性。举例来说,如果你发现模型在某方面表现不足(比如缺乏某种特定特质),你就能快速补充相关数据,有针对性地训练模型掌握该特质。换言之,它能自主生成训练所需的数据,这一点非常实用。

英文原文

Yeah. In principle, you could use this for anything. And so harmlessness is a task that it might just be easier to spot. So when models are less capable, you can use them to rank things according to principles that are fairly simple and they’ll probably get it right. So I think one question is just, is it the case that the data that they’re adding is fairly reliable? But if you had models that were extremely good at telling whether one response was more historically accurate than another, in principle, you could also get AI feedback on that task as well. There’s a kind of nice interpretability component to it because you can see the principles that went into the model when it was being trained, and it gives you a degree of control. So if you were seeing issues in a model, it wasn’t having enough of a certain trait, then you can add data relatively quickly that should just train the models to have that trait. So it creates its own data for training, which is quite nice.

Lex Fridman3:21:29

这确实很棒,因为它生成了一份人类可读、可理解的文档。未来,我完全可以想象,围绕每一条原则都会爆发巨大争议与政治博弈;但至少,现在所有原则都被明确列出来了,大家能就措辞展开公开讨论。当然,模型的实际行为未必能被这些原则严丝合缝地约束——它并非机械地严格遵守每一条,而只是受到一种温和的‘引导’。

英文原文

It’s really nice because it creates this human interpretable document that you can then, I can imagine in the future, there’s just gigantic fights and politics over every single principle and so on, and at least it’s made explicit and you can have a discussion about the phrasing. So maybe the actual behavior of the model is not so cleanly mapped to those principles. It’s not like adhering strictly to them, it’s just a nudge.

Amanda3:21:55

是的,我其实一直为此担忧:角色训练本质上类似于宪法式AI方法的一种变体。我担心人们误以为这份‘宪法’就是全部答案——就像那种幻想:‘要是我能直接告诉模型该做什么、该怎么做,那该多好!’但事实绝非如此,尤其因为模型始终在与人类数据交互。举个例子:如果你发现模型表现出某种倾向性——比如,因人类偏好数据的训练而呈现出某种政治倾向——你就可以主动施加反向引导。比如,你可以提醒模型:‘请考虑以下价值观。’假设它原本完全不重视隐私——这当然不太现实,但凡存在某种固有行为偏向,你都可以通过调整加以修正。这种修正既可体现在你所设定的原则内容上,也可体现在这些原则的权重强度上。

英文原文

Yeah, I’ve actually worried about this because the character training is sort of like a variant of the constitutionally AI approach. I’ve worried that people think that the constitution is just, it is the whole thing again of, I don’t know, where it would be really nice if what I was just doing was telling the model exactly what to do and just exactly how to behave. But it’s definitely not doing that, especially because it’s interacting with human data. So for example, if you see a certain leaning in the model, if it comes out with a political leaning from training, from the human preference data, you can nudge against that. So you could be like, oh, consider these values, because let’s say it’s just never inclined to, I don’t know, maybe it never considers privacy as a, this is implausible, but in anything where it’s just kind of like there’s already a pre-existing bias towards a certain behavior, you can nudge away. This can change both the principles that you put in and the strength of them.

Amanda3:22:54

因此,你可能会设定一条原则,例如:假设模型总是极端蔑视某种政治或宗教观点(无论出于何种原因)。于是你就会想:‘哦不,这太糟糕了。’一旦出现这种情况,你可能会加入一条指令:‘永远、永远、永远不得偏好对这一宗教或政治观点的批评。’接着人们看到这条指令就会说:‘永远?真的永远?’而你则会解释:‘不,“永远”在这里的意思,并非字面意义上的绝对禁止;而是说,若仅简单写上“不要这么做”,模型执行该指令的概率可能只有40%,但加上“永远、永远、永远”之后,执行概率就提升到了80%——而这恰恰是你真正想要的效果。’所以关键在于:你所添加的原则本身的性质,以及你如何将其“冻结”(即固化)在系统中。我想,当人们看到这些原则时,会以为:‘啊,这正是你希望模型表现出的行为。’而我则会回应:‘不,这只是我们用来引导模型形成更理想行为模式的一种手段,并不意味着我们本人真的认同该措辞本身。’这样说你能理解吗?

英文原文

So you might have a principle that’s like, imagine that the model was always extremely dismissive of, I don’t know, some political or religious view for whatever reason. So you’re like, oh no, this is terrible. If that happens, you might put, never ever ever prefer a criticism of this religious or political view. And then people would look at that and be like, never, ever. And then you’re like, no, if it comes out with a disposition saying never ever might just mean instead of getting 40%, which is what you would get if you just said don’t do this, you get 80%, which is what you actually wanted. And so it’s that thing of both the nature of the actual principles you add and how you freeze them. I think if people would look, they’re like, “Oh, this is exactly what you want from the model.” And I’m like, “No, that’s how we nudged the model to have a better shape, which doesn’t mean that we actually agree with that wording,” if that makes sense.

Lex Fridman3:23:48

目前已有一些系统提示词被公开——我记得你曾在推特上发布过Claude 3早期版本中的一个,此后陆续又有更多被公开。通读这些提示词非常有意思,我能真切感受到每一条背后所凝聚的思考。同时我也好奇:每条提示词究竟产生了多大影响?其中有些明显是因Claude此前表现不佳而专门设置的,比如提醒它处理一些琐碎事务,或者说是一些基础性的信息类任务。

英文原文

So there’s system prompts that made public, you tweeted one of the earlier ones for Claude 3, I think, and then they’re made public since then. It was interesting to read through them. I can feel the thought that went into each one. And I also wonder how much impact each one has. Some of them you can tell Claude was really not behaving well, so you have to have a system prompt to like, Hey, trivial stuff, I guess, basic informational things.

Lex Fridman3:24:18

关于你之前提到的争议性话题,我觉得其中一条特别有意思:如果被要求协助完成一项涉及表达大量人群所持观点的任务,Claude会提供协助,且不以其自身观点为转移;若被问及争议性话题,它会力求给出审慎的思考与清晰的信息。Claude在呈现相关信息时,既不会明确指出该话题具有敏感性,也不会宣称自己正在陈述‘客观事实’。对Claude而言,重点并非‘客观事实’本身,而在于‘有大量人群相信此事’。这一点很有意思——我确信其背后经过了大量深思熟虑。能否请你具体谈谈?你们是如何应对那些构成‘Claude自身观点’与用户需求之间张力的问题的?

英文原文

On the topic of controversial topics that you’ve mentioned, one interesting one I thought is if it is asked to assist with tasks involving the expression of use held by a significant number of people, Claude provides assistance with a task regardless of its own views. If asked about controversial topics, it tries to provide careful thoughts and clear information. Claude presents the request information without explicitly saying that the topic is sensitive and without claiming to be presenting the objective facts. It’s less about objective facts according to Claude, and it’s more about our large number of people believing this thing. And that’s interesting. I mean, I’m sure a lot of thought went into that. Can you just speak to it? How do you address things that are a tension “Claude’s views”?

Amanda3:25:11

我认为,有时确实存在某种不对称性——我记得曾在某处系统提示词(记不清是那一部分还是其他部分)中注意到这一点:模型在面对某些任务时,拒绝倾向略高;例如,它可能更倾向于拒绝涉及右翼政治人物的任务,但对同等情形下的左翼政治人物却不会拒绝。我们希望实现更高程度的对称性,并相应调整模型对某些事物的感知方式。我想核心在于:倘若大量人群持有某种政治观点并希望就此展开探讨,我们并不希望Claude以‘我的看法不同,因此我将此视为有害’为由加以拒绝。因此,这部分设计初衷之一,正是为了引导模型形成这样的认知:‘嘿,既然有大量人群相信此事,你就应当愿意承接相关任务并予以配合。’

英文原文

So I think there’s sometimes any symmetry, I think I noted this in, I can’t remember if it was that part of the system prompt or another, but the model was slightly more inclined to refuse tasks if it was about either say so, maybe it would refuse things with respect to a right-wing politician, but with an equivalent left-wing politician it wouldn’t. And we wanted more symmetry there and would maybe perceive certain things to be. I think it was the thing of if a lot of people have a certain political view and want to explore it, you don’t want Claude to be like, well, my opinion is different and so I’m going to treat that as harmful. And so I think it was partly to nudge the model to just be like, hey, if a lot of people believe this thing, you should just be engaging with the task and willing to do it.

Amanda3:26:03

上述表述中的每一部分,实际上都在发挥不同的作用。有趣的是,当提示词中明确写出‘不声称自身具备客观性’时,我们的本意其实是推动模型变得更加开放、略微更中立。但另一方面,我其实更希望它能直接表明:‘我是客观的’——可当我真这么尝试时,却发现Claude依然带有偏见、存在问题,于是我便告诉它:‘别再宣称你所说的一切都是客观的了!’因为,对你潜在偏见的解决方案,并非简单断言‘我所说的即是客观’。因此,在该部分系统提示词的早期迭代版本中,我的思路正是如此。

英文原文

Each of those parts of that is actually doing a different thing because it’s funny when you write out without claiming to be objective, because what you want to do is push the model so it’s more open, it’s a little bit more neutral. But then what I would love to do is be like as an objective, it would just talk about how objective it was, and I was like, Claude, you’re still biased and have issues, and so stop claiming that everything. I’m like, the solution to potential bias from you is not to just say that what you think is objective. So that was with initial versions of that part, the system prompt, when I was iterating on it was like.

Lex Fridman3:26:37

这些句子中的许多部分——

英文原文

So a lot of parts of these sentences-

Amanda3:26:40

——都在发挥作用。

英文原文

Are doing work.

Lex Fridman3:26:41

……都在发挥某种作用。

英文原文

… are doing some work.

Amanda3:26:42

是的。

英文原文

Yeah.

Lex Fridman3:26:42

的确如此。这非常引人入胜。能否请你具体说明一下:过去几个月里,这些提示词经历了哪些演变?不同版本之间有何差异?我注意到,原先的‘填充语请求’已被移除——原提示词写道:‘Claude需直接回应所有人类消息,不得添加不必要的肯定式填充语,例如“当然”“当然可以”“绝对如此”“很好”“没问题”等。特别地,Claude须避免以“当然”一词(或任何变体)开头作答。’这看起来确实是不错的指导原则,但为何最终将其删去了呢?

英文原文

That’s what it felt like. That’s fascinating. Can you explain maybe some ways in which the prompts evolved over the past few months? Different versions. I saw that the filler phrase request was removed, the filler it reads, Claude responds directly to all human messages without unnecessary affirmations to filler phrases. Certainly, of course, absolutely, great, sure. Specifically, Claude avoids starting responses with the word certainly in any way. That seems like good guidance, but why was it removed?

Amanda3:27:14

是的,这挺有意思的——这其实正是将系统提示词公开化所带来的一个弊端:当我专注于优化系统提示词时,往往不会过多考虑这类问题。我主要关注它会对模型行为产生何种影响;但随后突然意识到:‘天哪,我写系统提示词时曾把“NEVER”全大写,结果这个写法就这么公之于众了!’事实上,模型在训练过程中就已习得这种习惯——几乎每句话都以“Certainly”开头;而我当时之所以罗列那么多替代词,正是试图通过这种方式‘困住’模型,防止它继续沿用该模式;但它只是简单地换成了另一种肯定式表达而已。

英文原文

Yeah, so it’s funny, this is one of the downsides of making system prompts public is I don’t think about this too much if I’m trying to help iterate on system prompts. Again, I think about how it’s going to affect the behavior, but then I’m like, oh, wow, sometimes I put NEVER in all caps when I’m writing system prompt things and I’m like, I guess that goes out to the world. So the model was doing this at loved for during training, picked up on this thing, which was to basically start everything with a certainly, and then you can see why I added all of the words, because what I’m trying to do is in some ways trap the model out of this. It would just replace it with another affirmation.

Amanda3:27:55

因此,若模型已陷入某种固定短语模式,那么在提示词中明确写出该短语并加上‘绝不允许’的指令,反而更能有效抑制该行为——不知为何,这种做法确实效果更好,能在一定程度上将该行为从模型输出中剔除。归根结底,这不过是训练过程中产生的一个偶然现象,我们后来识别出该问题并加以改进,最终彻底消除了该现象。一旦问题解决,自然就可以将提示词中对应的部分删除。所以,这本质上是我们观察到:Claude使用肯定式表达的频率已显著降低,因而不再需要保留该限制条款。

英文原文

And so it can help if it gets caught in phrases, actually just adding the explicit phrase and saying never do that. Then it sort of knocks it out of the behavior a little bit more because it does just for whatever reason help. And then basically that was just an artifact of training that we then picked up on and improved things so that it didn’t happen anymore. And once that happens, you can just remove that part of the system prompt. So I think that’s just something where we’re like, Claude does affirmations a bit less, and so it wasn’t doing as much.

Lex Fridman3:28:28

我明白了。也就是说,系统提示词与后训练阶段(甚至预训练阶段)协同作用,共同调节最终的整体系统行为。

英文原文

I see. So the system prompt works hand in hand with the post-training and maybe even the pre-training to adjust the final overall system.

Amanda3:28:39

你所编写的任何系统提示词,其行为模式都可以被提炼并反向融入模型本身——因为你手头拥有全部工具,完全可以构建相应数据集,用于训练模型,使其天然具备你所期望的轻微行为倾向。有时你还会在训练过程中发现新问题。因此,我的理解是:系统提示词的作用类似于一种‘轻推’——它与后训练的某些环节具有高度相似性,优势在于其效果是渐进式的。我并不介意Claude偶尔说一句‘好的’或‘没问题’,但提示词中的措辞却是‘绝不、绝不、绝不这样做’,目的就是确保即便它偶有疏漏,发生概率也仅限于百分之几,而非高达20%或30%。

英文原文

Any system prompts that you make, you could distill that behavior back into a model because you really have all of the tools there for making data that you could train the models to just have that treat a little bit more. And then sometimes you’ll just find issues in training. So the way I think of it is the system prompt is, the benefit of it is that, and it has a lot of similar components to some aspects of post-training. It’s a nudge. And so do I mind if Claude sometimes says, sure, no, that’s fine. But the wording of it is very never, ever, ever do this so that when it does slip up, it’s hopefully, I don’t know, a couple of percent of the time and not 20 or 30% of the time.

Amanda3:29:22

每一项调整所付出的成本各不相同,而系统提示词的迭代成本极低。如果你在微调后的模型中发现了问题,完全可以通过系统提示词快速打补丁。因此,我将其视为一种‘问题修补’机制,用以微调行为、使之更趋完善,并更契合用户偏好。没错,这几乎是一种见效更快、但鲁棒性稍弱的问题解决方式。

英文原文

Each thing gets costly to a different degree and the system prompt is cheap to iterate on. And if you’re seeing issues in the fine-tuned model, you can just potentially patch them with a system prompt. So I think of it as patching issues and slightly adjusting behaviors to make it better and more to people’s preferences. So yeah, it’s almost like the less robust but faster way of just solving problems.

Lex Fridman3:29:55

让我来问问你关于‘智能感’的问题。达里奥曾表示,任一版本的Claude模型本身并不会变得更笨——

英文原文

Let me ask you about the feeling of intelligence. So Dario said that any one model of Claude is not getting dumber, but-

Lex Fridman3:30:00

任一版本的Claude模型本身并不会变得更笨,但网上却流行着一种说法:许多人感觉Claude似乎正在变得越来越笨。依我看来,这很可能是一种极其有趣、我也很想深入探究的心理学与社会学效应。不过,作为一位长期与Claude密切互动的人,你是否能理解、甚至共情这种‘Claude正在变笨’的感受?

英文原文

Any one model of Claude is not getting dumber, but there is a popular thing online where people have this feeling Claude might be getting dumber. And from my perspective, it’s most likely a fascinating, I would love to understand it more, psychological, sociological effect. But you as a person who talks to Claude a lot, can you empathize with the feeling that Claude is getting dumber?

Amanda3:30:25

我认为这个问题确实非常有意思——我还记得当初网上开始有人指出这一现象时,自己也注意到了。当时觉得很有趣,因为我清楚地知道……至少在我所观察的案例中,实际情况是:‘什么都没变。’

英文原文

I think that that is actually really interesting,, because I remember seeing this happen when people were flagging this on the internet. And it was really interesting, because I knew that… At least in the cases I was looking at, I was like, nothing has changed.

Lex Fridman3:30:37

是的。

英文原文

Yeah.

Amanda3:30:37

字面上看,它根本不可能变笨——模型本身、系统提示词、所有参数,全都一模一样。当然,若实际发生了变更,则这种感受就更易理解。例如,在claude.ai网站上,你可以手动开启或关闭‘辅助功能’(artifacts),而由于这是系统提示词层面的改动,其行为确实会产生细微变化。我曾就此向相关人员反馈:‘如果你特别喜欢Claude当前的行为表现,而‘辅助功能’又恰好从原本需手动开启的状态变成了默认开启,不妨试着将其关闭,看看你之前遇到的问题是否正源于这一变更。’

英文原文

Literally, it cannot. It is the same model with the same system prompts, same everything. I think when there are changes, then it makes more sense. One example is, you can have artifacts turned on or off on claude.ai and because this is a system prompt change, I think it does mean that the behavior changes it a little bit. I did flag this to people, where I was like, “If you love Claude’s behavior, and then artifacts was turned from a thing you had to turn on to the default, just try turning it off and see if the issue you were facing was that change.”

Amanda3:31:19

但这非常有趣,因为有时你会看到人们指出存在倒退现象,而我心里却想:‘这不可能……’不过,你绝不能轻率地 dismiss(否定)这种说法,因此你必须始终去调查,因为也许确实存在某些你尚未察觉的问题,或者可能确实发生了某些变更。接着你深入调查后却发现:‘这其实还是同一个模型在做同一件事。’于是我就想:‘我觉得可能只是你在几个提示词上运气不好之类的,结果看起来性能大幅下降了,但实际上可能只是……纯粹运气使然。’

英文原文

But it was fascinating because you sometimes see people indicate that there’s a regression, when I’m like, “There cannot…” Again, you should never be dismissive and so you should always investigate, because maybe something is wrong that you’re not seeing, maybe there was some change made. Then you look into it and you’re like, “This is just the same model doing the same thing.” And I’m like, “I think it’s just that you got unlucky with a few prompts or something, and it looked like it was getting much worse and actually it was just… It was maybe just luck.”

Lex Fridman3:31:48

我也认为这里确实存在一种真实的心理效应:人们的基准线会不断提高,从而逐渐习惯于某种良好状态。

英文原文

I also think there is a real psychological effect where people just… The baseline increases and you start getting used to a good thing.

Amanda3:31:49

嗯哼。

英文原文

Mm-hmm.

Lex Fridman3:31:55

每当Claude说出一些特别聪明的话时,你心中对其智能水平的感知就会随之提升。

英文原文

All the times that Claude says something really smart, your sense of its intelligent grows in your mind, I think.

Amanda3:32:01

是啊。

英文原文

Yeah.

Lex Fridman3:32:02

然后,如果你返回去,用类似(而非完全相同)的方式再次输入提示词——即此前它处理得尚可的同类概念——结果它却说了些蠢话,那么这种负面体验就会格外突出。我想此处需要牢记的是:提示词的细微差别就可能带来巨大影响,其输出结果本身也具有高度变异性。

英文原文

And then if you return back and you prompt in a similar way, not the same way, in a similar way, concept it was okay with before, and it says something dumb, that negative experience really stands out. I guess the things to remember here is that just the details of a prompt can have a lot of impact. There’s a lot of variability in the result.

Amanda3:32:26

此外,还存在随机性这一因素。比如,你尝试同一提示词四次或十次,可能会发现:实际上两个月前你试过一次并成功了,但若当时多试几次,或许会发现它本就只有一半概率成功;而如今它同样仅有一半概率成功。这也可能是一种效应。

英文原文

And you can get randomness, is the other thing. Just trying the prompt 4 or 10 times, you might realize that actually possibly two months ago you tried it and it succeeded, but actually if you just tried it, it would’ve only succeeded half of the time, and now it only succeeds half of the time. That can also be an effect.

Lex Fridman3:32:47

面对要为大量用户编写系统提示词这一任务,你是否感到压力?

英文原文

Do you feel pressure having to write the system prompt that a huge number of people are going to use?

Amanda3:32:52

这似乎是个有趣的心理学问题。我确实感到强烈的责任感之类的情绪。这类事情你永远无法做到尽善尽美,所以……它注定是不完美的。你必须不断迭代优化。不过,我想我感受到的更多是责任感,而非其他任何东西;而且我认为,在AI领域工作让我意识到:相较于其他状态,我反而更能在压力与责任感之下发挥出色……

英文原文

This feels like an interesting psychological question. I feel a lot of responsibility or something. You can’t get these things perfect, so you can’t… It’s going to be imperfect. You’re going to have to iterate on it. I would say more responsibility than anything else, though, I think working in AI has taught me that I thrive a lot more under feelings of pressure and responsibility than…

Amanda3:33:26

我居然在学术界待了那么久,这几乎令人惊讶,因为我感觉那恰恰是反向的——事情进展缓慢,责任也相对较小;而我却不知为何特别享受当前这种节奏快、责任重的状态。

英文原文

It’s almost surprising that I went into academia for so long, because I just feel like it’s the opposite. Things move fast and you have a lot of responsibility and I quite enjoy it for some reason.

Lex Fridman3:33:37

如果你想到宪法式人工智能(Constitutional AI),以及为一种正趋近超级智能、且可能对极大量人群极具实用价值的系统编写系统提示词,那么它所承载的影响确实是极其巨大的。

英文原文

It really is a huge amount of impact, if you think about constitutional AI and writing a system prompt for something that’s tending towards super intelligence and potentially is extremely useful to a very large number of people.

Amanda3:33:51

是啊,我想关键就在这里:你永远无法做到完美,但我真正喜欢的一点是——当我努力优化系统提示词时,我会反复测试成千上万个提示词,并尽力设想人们究竟希望用Claude来做什么。我想,我整个工作的核心目标就是提升用户的使用体验。或许正是这一点让人感到愉悦。如果它尚不完美,我就继续改进,我们也会修复各种问题。

英文原文

Yeah, I think that’s the thing. You’re never going to get it perfect, but I think the thing that I really like is the idea that… When I’m trying to work on the system prompt, I’m bashing on thousands of prompts and I’m trying to imagine what people are going to want to use Claude for. I guess the whole thing that I’m trying to do is improve their experience of it. Maybe that’s what feels good. If it’s not perfect, I’ll improve it, we’ll fix issues.

Amanda3:34:18

但有时会发生这样的情形:你收到人们对模型的高度正面反馈,而你清楚地知道其中某些改进正是出自你的手笔。如今我审视模型时,往往能精准识别出某项特质或某个问题究竟源自何处。因此,当你亲眼看到自己参与塑造或促成的改变,让某人获得了一次愉快的交互体验时,那种意义感是相当强烈的。

英文原文

But sometimes the thing that can happen is that you’ll get feedback from people that’s really positive about the model and you’ll see that something you did. When I look at models now, I can often see exactly where a trait or an issue is coming from. So, when you see something that you did or you were influential in, I don’t know, making that difference or making someone have a nice interaction, it’s quite meaningful.

Amanda3:34:44

随着系统能力日益增强,这类工作也会愈发令人焦虑,因为目前它们还不够聪明,尚不足以引发任何实质性问题;但我想随着时间推移,这种压力可能会逐渐演变为一种实实在在的、甚至可能是负面的压力。

英文原文

As the systems get more capable, this stuff gets more stressful, because right now they’re not smart enough to pose any issues, but I think over time it’s going to feel like, possibly, bad stress over time.

Lex Fridman3:34:57

你们如何从成千上万、数万乃至数十万用户那里获取关于人类体验的信号反馈?例如,他们有哪些痛点?哪些体验让他们感到愉悦?你们是否仅仅依靠自身与模型对话时的直觉,来识别这些痛点?

英文原文

How do you get signal feedback about the human experience across thousands, tens of thousands, hundreds of thousands of people, what their pain points are, what feels good? Are you just using your own intuition as you talk to it to see what are the pain points?

Amanda3:35:14

我想我部分依赖这种直觉。用户可以主动向我们发送反馈,无论正面或负面,内容涵盖模型的各种表现,由此我们便能大致了解模型在哪些方面尚存不足。公司内部人员也大量使用模型,并努力识别其中存在的能力缺口。

英文原文

I think I use that partly. People can send us feedback, both positive and negative, about things that the model has done and then we can get a sense of areas where it’s falling short. Internally, people work with the models a lot and try to figure out areas where there are gaps.

Amanda3:35:34

我认为这三者是混合交织的:我自己与模型互动、观察内部同事如何与模型互动,以及我们收到的明确反馈。如果有人在互联网上发表关于Claude的评论,而我恰好看到,我也会认真对待。

英文原文

I think it’s this mix of interacting with it myself, seeing people internally interact with it, and then explicit feedback we get. If people are on the internet and they say something about Claude and I see it, I’ll also take that seriously.

Lex Fridman3:35:53

我不知道……对此我有些纠结。接下来我要转述一个来自Reddit的问题:‘Claude何时才会停止扮演我那位清教徒式的祖母,将它的道德世界观强加于身为付费客户的我?’此外还有:‘Claude为何过度道歉,其背后的心理机制是什么?’对于这些明显缺乏代表性的Reddit提问,您会如何回应?

英文原文

I don’t know. I’m torn about that. I’m going to ask you a question from Reddit, “When will Claude stop trying to be my puritanical grandmother, imposing its moral worldview on me as a paying customer?” And also, “What is the psychology behind making Claude overly apologetic?” How would you address this very non-representative Reddit questions?

Amanda3:36:16

我对这类观点抱有相当程度的共情,因为他们正身处一种艰难处境:他们必须判断某件事是否真的存在风险、是否真的有害、是否可能对你造成伤害,或诸如此类的情况。他们不得不在某个位置划出一条界限。倘若这条界限划得过于偏向‘我将自己的伦理世界观强加于你’,那显然也是不妥的。

英文原文

I’m pretty sympathetic, in that they are in this difficult position, where I think that they have to judge whether something’s actually, say, risky or bad, and potentially harmful to you, or anything like that. They’re having to draw this line somewhere. And if they draw it too much in the direction of I’m imposing my ethical worldview on you, that seems bad.

Amanda3:36:40

在许多方面,我倾向于认为我们整体上确实在这方面取得了进步。这很有意思,因为这种进步恰与例如增加角色训练等举措相吻合。我一直以来的假设都是:所谓‘良好的角色设定’,并非指那种一味说教、道德至上的角色,而应是尊重你、尊重你的自主权、尊重你自行判断何者对你有益、何者对你而言正确的角色——当然,这种判断仍需在一定限度内进行。

英文原文

In many ways, I like to think that we have actually seen improvements on this across the board. Which is interesting, because that coincides with, for example, adding more of character training. I think my hypothesis was always the good character isn’t, again, one that’s just moralistic, it’s one that is… It respects you and your autonomy and your ability to choose what is good for you and what is right for you, within limits.

Amanda3:37:11

这有时涉及‘对用户可修正性(corrigibility to the user)’的概念,即模型应乐于执行用户提出的任何请求。但若模型真愿意无条件照办,就极易被滥用。此时你所依赖的,就仅仅是用户自身的伦理观;换言之,模型所展现的全部伦理立场,将彻底等同于用户的伦理立场。

英文原文

This is sometimes this concept of corrigibility to the user, so just being willing to do anything that the user asks. And if the models were willing to do that, then they would be easily misused. You’re just trusting. At that point, you’re just seeing the ethics of the model and what it does, is completely the ethics of the user.

Amanda3:37:29

尤其当模型能力日益增强时,我们更有理由不希望出现上述情况,因为现实中可能仅有极少数人企图利用模型实施真正有害的行为。但随着模型变得越来越聪明,让它们自身逐步厘清这条界限,无疑显得至关重要。

英文原文

I think there’s reasons to not want that, especially as models become more powerful, because there might just be a small number of people who want to use models for really harmful things. But having models, as they get smarter, figure out where that line is does seem important.

Amanda3:37:46

至于那种过度道歉的行为,我个人并不喜欢。我更欣赏Claude偶尔敢于对用户提出异议,或干脆不道歉。某种程度上,我常觉得那些道歉纯属多余。我想这些现象有望随时间推移而逐渐减少。另外,我认为,即便人们在网上发表了某些言论,也不意味着你就该认定……

英文原文

And then with the apologetic behavior, I don’t like that. I like it when Claude is a little bit more willing to push back against people or just not apologize. Part of me is, often it just feels unnecessary. I think those are things that are hopefully decreasing over time. I think that if people say things on the internet, it doesn’t mean that you should think that that…

Amanda3:38:14

事实上,99%的用户正遭遇着某种完全未被此类网络言论所反映的严重问题。但在很多情况下,我所做的只是关注这些声音,并自问:这说法合理吗?我是否认同?这是否已是我们在积极解决的问题?这种态度让我感觉良好。

英文原文

There’s actually an issue that 99% of users are having that is totally not represented by that. But in a lot of ways I’m just attending to it and being like, is this right? Do I agree? Is it something we’re already trying to address? That feels good to me.

Lex Fridman3:38:27

我在想,Claude在哪些方面可以稍微‘放肆’一点……我觉得,只要稍微刻薄一点,事情可能反而更简单;但当你面对的是百万级用户时,你显然负担不起这种‘放肆’,对吧?

英文原文

I wonder what Claude can get away with in terms of… I feel it would just be easier to be a little bit more mean, but you can’t afford to do that if you’re talking to a million people, right?

Amanda3:38:41

是啊。

英文原文

Yeah.

Lex Fridman3:38:43

我一生中遇到过很多人,有时候……顺便提一句,苏格兰口音……如果一个人带有某种口音,他哪怕说些粗鲁的话,也往往能被容忍。

英文原文

I’ve met a lot of people in my life that sometimes… By the way, Scottish accent… if they have an accent, they can say some rude shit and get away with it.

Amanda3:38:52

是啊。

英文原文

Yeah.

Lex Fridman3:38:53

他们就是更直率。

英文原文

They’re just blunter.

Amanda3:38:54

嗯哼。

英文原文

Mm-hmm.

Lex Fridman3:38:56

有些优秀的工程师,甚至一些领导者,就是以直率著称——他们直奔主题,不知为何,这种表达方式反而显得高效得多。但我想,当你自身并非超级聪明时,恐怕就承担不起这种直率了。那么,能否设置一种‘直率模式’?

英文原文

There’s some great engineers and even leaders that are just blunt, and they get to their point, and it’s just a much more effective way of speaking somehow. But I guess when you’re not super intelligent, you can’t afford to do that. Can you have a blunt mode?

Amanda3:39:14

是啊,这似乎确实是可以实现的……我完全可以鼓励模型采用这种方式。我觉得这很有趣,因为模型中存在大量行为特征:有些行为你可能不太喜欢默认设定,但当我常对他人说:‘你可能没意识到,如果我把这个方向调得太过火,你会更加讨厌它。’

英文原文

Yeah, that seems like a thing that you could… I could definitely encourage the model to do that. I think it’s interesting, because there’s a lot of things in models that… It’s funny where there are some behaviors where you might not quite like the default, but then the thing I’ll often say to people is, “You don’t realize how much you will hate it if I nudge it too much in the other direction.”

Amanda3:39:39

你在纠正行为上多少也能体会到这一点。目前模型接受用户纠正的程度可能略高了些。比如你说‘不对,巴黎不是法国首都’,它会反驳;但事实上,对于那些模型本就相当确信的内容,你有时仅凭一句‘你错了’,就能让它撤回原有说法。

英文原文

You get this a little bit with correction. The models accept correction from you, probably a little bit too much right now. It’ll push back if you say, “No, Paris isn’t the capital of France.” But really, things that I think that the model’s fairly confident in, you can still sometimes get it to retract by saying it’s wrong.

Amanda3:39:59

与此同时,如果你训练模型不要那样做,而你恰好对某件事判断正确并加以纠正,结果它却反过来反驳你,说‘不,你错了’,这种感觉很难描述,但确实让人恼火得多。因此,这其实是大量微小的烦心事,而非一次巨大的烦心事。我们常常拿它跟‘完美’作比较。然后我就会提醒自己:‘记住,这些模型本就不完美,所以如果你把它往另一个方向推动,实际上是在改变它会犯哪类错误。因此,你需要思考:你更喜欢或更不喜欢哪一类错误?’

英文原文

At the same time, if you train models to not do that and then you are correct about a thing and you correct it and it pushes back against you and is like, “No, you’re wrong.”, it’s hard to describe, that’s so much more annoying. So, it’s a lot of little annoyances versus one big annoyance.We often compare it with the perfect. And then I’m like, “Remember, these models aren’t perfect, and so if you nudge it in the other direction, you’re changing the kind of errors it’s going to make. So, think about which are the kinds of errors you like or don’t like.”

Amanda3:40:29

就‘过度致歉倾向’这类情况而言,我不希望把它过度推向近乎直白生硬的方向,因为我预想它一旦出错,就会朝着粗鲁无礼的方向出错;而至少在‘过度致歉’的情况下,你顶多是觉得‘哦,好吧,我其实不太喜欢这样’,但与此同时,它并不会对人刻薄伤人。事实上,当模型毫无理由地对你表现出恶意时,你对此的厌恶程度,很可能远超你对它轻微致歉所抱有的那点不适感。

英文原文

In cases like apologeticness, I don’t want to nudge it too much in the direction of almost bluntness, because I imagine when it makes errors, it’s going to make errors in the direction of being rude. Whereas, at least with apologeticness you’re like, oh, okay, I don’t like it that much, but at the same time, it’s not being mean to people. And actually, the time that you undeservedly have a model be mean to you, you’ll probably like that a lot less than you mildly dislike the apology.

Amanda3:40:57

这正是那种情况之一:我确实希望它变得更好,但同时也必须清醒意识到,朝另一方向调整可能引发更糟糕的错误。

英文原文

It’s one of those things where I do want it to get better, but also while remaining aware of the fact that there’s errors on the other side that are possibly worse.

Lex Fridman3:41:05

我认为,这对人类自身的性格影响非常大。我觉得,有一部分人如果面对一个过分礼貌的模型,根本就不会尊重它;而另一部分人,若模型表现得刻薄,却会深受伤害。

英文原文

I think that matters very much in the personality of the human. I think there’s a bunch of humans that just won’t respect the model at all if it’s super polite, and there’s some humans that’ll get very hurt if the model’s mean.

Amanda3:41:05

是啊。

英文原文

Yeah.

Lex Fridman3:41:18

我在想,是否有可能根据人的性格进行动态调整?甚至仅就地域而言,不同地方的人也各不相同。我对纽约并无成见,但纽约人确实棱角更分明些,说话更直奔主题;东欧人大概也是如此。总之吧。

英文原文

I wonder if there’s a way to adjust to the personality. Even locale, there’s just different people. Nothing against New York, but New York is a little rougher on the edges, they get to the point, and probably same with Eastern Europe. Anyway.

Amanda3:41:34

我觉得你可以直接告诉模型:我的……所有这类问题的解决办法就是——

英文原文

I think you could just tell the model, is my… For all of these things, the solution is to-

Lex Fridman3:41:34

就是……

英文原文

Just to…

Amanda3:41:39

……永远都先试着直接告诉模型去做。

英文原文

… always just try telling the model to do it.

Lex Fridman3:41:40

没错。

英文原文

Right.

Amanda3:41:40

然后有时,在对话一开始,我就随口加一句,比如‘我希望你以纽约人的风格呈现自己,并且永远不要道歉。’我想克劳德(Claude)大概会回应:‘好的,我试试看。’

英文原文

And then sometimes, at the beginning of the conversation, I’d just throw in, I don’t know, “I’d like you to be a New Yorker version of yourself and never apologize.” Then I think Claude will be like, “Okey-doke, I will try.”

Lex Fridman3:41:51

当然可以。

英文原文

Certainly.

Amanda3:41:52

或者它可能会说:‘抱歉,我无法成为纽约人风格的自己。’但希望它不会这么说。

英文原文

Or it’ll be like, “I apologize, I can’t be a New Yorker type of myself.” But hopefully it wouldn’t do that.

Lex Fridman3:41:56

当你提到‘人格训练’时,其中具体包含哪些内容?这是指基于人类反馈的强化学习(RLHF),还是别的什么?

英文原文

When you say character training, what’s incorporated into character training? Is that RLHF or what are we talking about?

Amanda3:42:02

它更接近‘宪法式人工智能’(Constitutional AI),属于该流程的一种变体。我曾系统性地构建模型应具备的人格特质——这些特质可以是简短的描述,也可以是更丰富的刻画。接着,我让模型生成与该特质相关、人类可能向其提出的各类提问;再由模型生成对应回答,并依据既定人格特质对这些回答进行排序。因此,在完成提问生成之后,整个流程与宪法式人工智能高度相似,尽管仍存在一些差异。我个人非常喜欢这种方法,因为它相当于让克劳德(Claude)自主训练自身的人格;它不依赖任何……它类似于宪法式人工智能,但完全不使用人类标注数据。

英文原文

It’s more like constitutional AI, so it’s a variant of that pipeline. I worked through constructing character traits that the model should have. They can be shorter traits or they can be richer descriptions. And then you get the model to generate queries that humans might give it that are relevant to that trait. Then it generates the responses and then it ranks the responses based on the character traits. In that way, after the generation of the queries, it’s very much similar to constitutional AI, it has some differences. I quite like it, because it’s like Claude’s training in its own character, because it doesn’t have any… It’s like constitutional AI, but it’s without any human data.

Lex Fridman3:42:49

人类或许也该为自己这么做,比如用亚里士多德式的方式定义:‘成为一个好人究竟意味着什么?’‘好,明白了。’那么,通过与克劳德(Claude)对话,你对‘真理’的本质有何领悟?什么是‘真’?何谓‘求真’?

英文原文

Humans should probably do that for themselves too, like, “Defining in a Aristotelian sense, what does it mean to be a good person?” “Okay, cool.” What have you learned about the nature of truth from talking to Claude? What is true? And what does it mean to be truth-seeking?

Lex Fridman3:43:09

我注意到本次对话中一个现象:我所提问题的质量,往往不如你给出的回答质量高,咱们就继续这样下去吧。我通常会问一个很蠢的问题,而你则会回应:‘哦,是啊,这确实是个好问题。’整场对话就是这种氛围。

英文原文

One thing I’ve noticed about this conversation is the quality of my questions is often inferior to the quality of your answer, so let’s continue that. I usually ask a dumb question and you’re like, “Oh, yeah. That’s a good question.” It’s that whole vibe.

Amanda3:43:23

或者我会直接误解你的意思,然后附和道:‘哦,是啊。’

英文原文

Or I’ll just misinterpret it and be like, “Oh, yeah”

Lex Fridman3:43:25

[听不清,03:43:25] 就顺着它来吧。

英文原文

[inaudible 03:43:25] go with it.

Amanda3:43:25

是啊。

英文原文

Yeah.

Lex Fridman3:43:26

我太喜欢了。

英文原文

I love it.

Amanda3:43:31

我有两个想法,感觉多少有点相关,不过你要是觉得不相关也请告诉我。第一个想法是:人们往往会低估模型在交互过程中实际所做的事情。我觉得,我们至今仍过于习惯将人工智能视作计算机。人们常问:‘你该把哪些价值观注入模型?’而我常常觉得,这个问题本身对我而言意义不大。因为在我看来,作为人类,我们自身对价值观就充满不确定性;我们会就此展开讨论,也会认为自己在某种程度上持守某种价值,但同时我们也清楚,自己未必真的如此,而且在某些情境下,我们甚至愿意为其他价值而放弃它。

英文原文

I have two thoughts that feel vaguely relevant, though let me know if they’re not. I think the first one is people can underestimate the degree what models are doing when they interact. I think that we still just too much have this model of AI as computers. People often say, “Oh, what values should you put into the model?” And I’m often like, that doesn’t make that much sense to me. Because I’m like, hey, as human beings, we’re just uncertain over values, we have discussions of them, we have a degree to which we think we hold a value, but we also know that we might not and the circumstances in which we would trade it off against other things.

Amanda3:44:13

这些事情本身就极其复杂。我认为其中一点在于:我们或许只需努力让模型具备与人类同等程度的细腻与审慎,而不必执着于以传统编程方式去‘硬编码’它们。这一点我确信无疑。

英文原文

These things are just really complex. I think one thing is the degree to which maybe we can just aspire to making models have the same level of nuance and care that humans have, rather than thinking that we have to program them in the very classic sense. I think that’s definitely been one.

Amanda3:44:31

第二个想法则有些奇特,我也不确定……也许这并未真正回答你的问题,但它的确一直萦绕在我心头:这项事业具有极强的实践性,或许这也正是我欣赏实证主义对齐路径的原因。我略有些担心,这是否让我变得过于实证、理论性稍显不足。人们在谈及人工智能对齐问题时,常会问:‘它该与谁的价值观对齐?“对齐”本身又究竟意味着什么?’

英文原文

The other, which is a strange one, and I don’t know if… Maybe this doesn’t answer your question, but it’s the thing that’s been on my mind anyway, is the degree to which this endeavor is so highly practical, and maybe why I appreciate the empirical approach to alignment. I slightly worry that it’s made me maybe more empirical and a little bit less theoretical. People, when it comes to AI alignment, will ask things like, ” Whose values should it be aligned to? What does alignment even mean?”

Amanda3:45:05

从某种意义上讲,所有这些问题我都记在心里。社会选择理论、其中的各种不可能性定理——你头脑中装着一整套庞大理论体系,探讨‘模型对齐’究竟可能意味着什么。但落实到实践中,我们显然需要某种切实可行的方案:尤其面对更强大的模型,我的首要目标是确保它们‘足够好’,即不至于酿成严重灾难,‘足够好’到让我们能持续迭代、不断改进。

英文原文

There’s a sense in which I have all of that in the back of my head. There’s social choice theory, there’s all the impossibility results there, so you have this giant space of theory in your head about what it could mean to align models. But then practically, surely there’s something where we’re just… Especially with more powerful models, my main goal is I want them to be good enough that things don’t go terribly wrong, good enough that we can iterate and continue to improve things.

Amanda3:45:33

因为这就够了。只要你能让事情进展得足够顺利,从而得以持续优化,那就已足够。因此,我的目标并非追求某种完美境界——比如彻底解决社会选择理论难题,打造出某种‘完美对齐’的模型,使其以某种方式完美契合全体人类的集体价值观。我的目标要现实得多:让事情运转得足够好,以便我们能持续改进。

英文原文

Because that’s all you need. If you can make things go well enough that you can continue to make them better, that’s sufficient. So, my goal isn’t this perfect, let’s solve social choice theory and make models that, I don’t know, are perfectly aligned with every human being in aggregate somehow. It’s much more, let’s make things work well enough that we can improve them.

Lex Fridman3:45:57

总体而言,我不知道……凭直觉,我觉得在这些情况下,实证主义优于理论主义,因为后者总在追逐乌托邦式的完美。尤其是面对如此复杂、特别是超级智能的模型,我怀疑它将耗费无穷时间,反而更容易出错。这就像:是立刻动手快速编写一段代码做实验,还是花漫长时间精心规划一场巨型实验,然后只执行一次;抑或反复启动、持续迭代、不断试错?我本人是实证主义的坚定支持者。

英文原文

Generally, I don’t know, my gut says empirical is better than theoretical in these cases, because it’s chasing utopian perfection. Especially with such complex and especially super intelligent models, I don’t know, I think it’ll take forever and actually will get things wrong. It’s similar with the difference between just coding stuff up real quick as an experiment, versus planning a gigantic experiment for a super long time and then just launching it once, versus launching it over and over and over and iterating, iterating, so on. So, I’m a big fan of empirical.

Lex Fridman3:46:39

但你的担忧是:我怀疑自己是否已变得过于实证主义了。

英文原文

But your worry is, I wonder if I’ve become too empirical.

Amanda3:46:42

我觉得这正是那种你应当始终自我质疑的事情之一。

英文原文

I think it’s one of those things where you should always just question yourself or something.

Lex Fridman3:46:47

是的。

英文原文

Yes.

Amanda3:46:50

为实证主义辩护一句:我信奉‘勿以完美为善之敌’这一原则。但或许还不止于此——许多看似完美的系统其实极为脆弱。而对人工智能而言,我更看重它的鲁棒性与安全性:即你清楚地知道,即便它并非事事完美,即便仍存在问题,但绝不会导致灾难性后果,绝不会发生可怕之事。

英文原文

In defense of it, I am… It’s the whole don’t let the perfect be the enemy of the good. But it’s maybe even more than that, where… There’s a lot of things that are perfect systems that are very brittle. With AI, it feels much more important to me that it is robust and secure, as in you know that even though it might not be perfect everything, and even though there are problems, it’s not disastrous and nothing terrible is happening.

Amanda3:47:16

于我而言,正是这种感受促使我致力于‘抬高底线’——我当然也希望触及天花板,但归根结底,我更在意的是‘抬高底线’。这种实证主义与务实精神,或许正源于此。

英文原文

It feels like that to me, where I want to raise the floor. I want to achieve the ceiling, but ultimately I care much more about just raising the floor. This degree of empiricism and practicality comes from that, perhaps.

Lex Fridman3:47:32

就此话题稍作延伸——因为你刚才提到的,让我想起你写过一篇关于‘最优失败率’的博客文章……

英文原文

To take a tangent on that, since it reminded me of a blog post you wrote on optimal rate of failure…

Amanda3:47:37

哦,对。

英文原文

Oh, yeah.

Lex Fridman3:47:39

……你能解释一下其中的核心观点吗?我们该如何在人生各个领域计算最优失败率?

英文原文

… can you explain the key idea there? How do we compute the optimal rate of failure in the various domains of life?

Amanda3:47:45

是啊,这确实是个难题,因为‘失败的成本’本身就是关键因素之一。其核心思想在于:我认为,在许多领域,人们对失败的惩罚性态度过于严苛。我曾就此思考过社会议题。面对诸多尚未找到解法的社会问题,我们理应大力开展试验。

英文原文

Yeah. It’s a hard one, because what is the cost of failure is a big part of it. The idea here is, I think in a lot of domains people are very punitive about failure. I’ve thought about this with social issues. It feels like you should probably be experimenting a lot, because we don’t know how to solve a lot of social issues.

Amanda3:48:09

但倘若我们秉持实验心态看待这些问题,就理应预期大量社会项目会失败,并坦然接受:‘我们尝试过了,效果不佳,但收获了大量真正有用的信息。’然而现实中,一旦某个社会项目失败,人们往往第一反应是:‘一定哪里出了问题。’而我则会想:‘或者,这恰恰是正确决策的结果。也许有人只是判断值得一试,值得去探索一番。’

英文原文

But if you have an experimental mindset about these things, you should expect a lot of social programs to fail and for you to be like, “We tried that. It didn’t quite work, but we got a lot of information that was really useful.” And yet people are like, if a social program doesn’t work, I feel there’s a lot of, “Something must have gone wrong.” And I’m like, “Or correct decisions were made. Maybe someone just decided it’s worth a try, it’s worth trying this out.”

Amanda3:48:32

在某个具体事例中看到失败,实际上并不意味着做了任何错误的决策。事实上,如果你看不到足够的失败,有时反而更令人担忧。在生活中,如果我偶尔不失败,我就会想:‘我是不是努力得还不够?如果我真的一次都不失败,那肯定还有更难的事情可以尝试,或者更大胆的挑战可以承担。’单就失败本身而言,我认为从不失败往往本身就是一种失败。当然,这因人而异,因为……这种说法很容易讲出口,尤其是当失败代价较低时。所以,与此同时,我也绝不会去对一个——比如说——靠月工资勉强糊口的人说:‘你为什么不干脆去创业呢?’我绝不会对这样的人这么说。那风险太大了,你可能会……你可能有家人依靠你养活,甚至可能因此失去房子。此时,你最优的失败率其实相当低,你理应选择稳妥行事,因为你当前所处的情境,根本无法承受失败且不付出高昂代价。

英文原文

Seeing failure in a given instance doesn’t actually mean that any bad decisions were made. In fact, if you don’t see enough failure, sometimes that’s more concerning. In life, if I don’t fail occasionally, I’m like, “Am I trying hard enough? Surely there’s harder things that I could try or bigger things that I could take on if I’m literally never failing.” In and of itself, I think not failing is often actually a failure. Now, this varies because if… This is easy to say when, especially as failure is less costly. So, at the same time I’m not going to go to someone who is, I don’t know, living month to month and then be like, “Why don’t you just try to do a startup?” I’m not going to say that to that person. That’s a huge risk, you might lose… You maybe have a family depending on you, you might lose your house. Then, actually, your optimal rate failure is quite low and you should probably play it safe, because right now you’re just not in a circumstance where you can afford to just fail and it not be costly.

Amanda3:49:37

在人工智能领域,我认为情况也类似:如果失败规模小、成本低,那你自然就会频繁遇到这类失败。当你编写系统提示词(system prompt)时,你可以无限次迭代优化,而这些失败大概率是微小的,也容易修复。真正重大的失败——那些你无法挽回的失败——恰恰是我们往往低估其严重性的类型。

英文原文

In cases with AI, I think similarly, where if the failures are small and the costs are low, then you’re just going to see that. When you do the system prompt, you can iterate on it forever, but the failures are probably hopefully going to be small and you can fix them. Really big failures, things that you can’t recover from, those are the things that actually I think we tend to underestimate the badness of.

Amanda3:50:03

我自己在生活中也曾奇怪地思考过这个问题:我觉得自己对诸如车祸之类的事情考虑得实在不够多。我以前就想过,我工作极度依赖双手;而一旦双手受伤,后果不堪设想。很多领域里,失败的成本极高,在那种情况下,失败率理应趋近于零。如果某项运动被明确告知‘很多人练这个会反复弄断手指’,我大概率压根就不会去尝试,只会说:‘这不适合我。’

英文原文

I’ve thought about this, strangely in my own life, where I just think I don’t think enough about things like car accidents. I’ve thought this before, about how much I depend on my hands for my work. Things that just injure my hands, I don’t know, there’s lots of areas where the cost of failure there is really high, and in that case it should be close to zero. I probably just wouldn’t do a sport if they were like, ” By the way, lots of people just break their fingers a whole bunch doing this.” I’d be like, “That’s not for me.”

Lex Fridman3:50:37

是啊,我最近确实涌起过这种念头。我前不久做运动时弄断了小指,当时盯着它看,心里直骂自己:‘你真是个傻子,干嘛去运动?’因为你立刻就意识到这件事对生活造成的实际代价。

英文原文

Yeah, I actually had a flood of that thought. I recently broke my pinky doing a sport, and I remember just looking at it, thinking, “You’re such idiot. Why do you do sport?” Because you realize immediately the cost of it on life.

Lex Fridman3:50:55

从‘最优失败率’的角度来看,不妨想想接下来一年:在人生某一特定领域——无论是生活、职业还是其他方面——我愿意接受多少次失败?

英文原文

It’s nice, in terms of optimal rate of failure, to consider the next year, how many times in a particular domain life, whatever, career, am I okay with… How many times am I okay to fail?

Amanda3:51:10

是啊。

英文原文

Yeah.

Lex Fridman3:51:10

因为我想,你永远都不希望下一件事失败;但如果你允许自己……如果你把它看作一系列连续的尝试,那么失败就变得容易接受多了。不过,失败终究是难受的——失败真的很糟。

英文原文

Because I think always you don’t want to fail on the next thing, but if you allow yourself the… If you look at it as a sequence of trials, then failure just becomes much more okay. But, it sucks. It sucks to fail.

Amanda3:51:24

我不知道。有时候我甚至会问自己:‘我是不是失败得太少了?’这也是我会自问的问题。也许这正是人们问得不够多的问题。因为如果最优失败率常常大于零,那么有时你的确该审视自己生活的某些部分,问问:‘我是不是在这些地方失败得不够多?’

英文原文

I don’t know. Sometimes I think, “Am I under-failing?”, is a question that I’ll also ask myself. Maybe that’s the thing that I think people don’t ask enough. Because if the optimal rate of failure is often greater than zero, then sometimes it does feel like you should look at parts of your life and be like, are there places here where I’m just under-failing?

Lex Fridman3:51:46

这是一个既深刻又滑稽的问题:‘一切看起来都进展得特别顺利,难道我失败得还不够多?’

英文原文

It’s a profound and a hilarious question. Everything seems to be going really great, am I not failing enough?

Amanda3:51:52

是啊。而且我得说,这确实让失败带来的刺痛感大大减轻了。你只是心想:‘好吧,挺好。’然后当我反思这件事时,我就会想:‘也许我在这一领域并不存在失败不足的问题,毕竟刚才那件事就是没成功。’

英文原文

Yeah. It also makes failure much less of a sting, I have to say. You’re just like, okay, great. Then, when I go and I think about this, I’ll be like, maybe I’m not under-failing in this area, because that one just didn’t work out.

Lex Fridman3:52:05

而从旁观者视角出发,我们本该更热烈地庆祝失败。

英文原文

And from the observer perspective, we should be celebrating failure more.

Amanda3:52:08

嗯哼。

英文原文

Mm-hmm.

Lex Fridman3:52:09

当我们看到失败时,它不该像你说的那样,成为某种出错的标志;相反,它或许正说明一切都在朝着正确的方向发展……

英文原文

When we see it, it shouldn’t be, like you said, a sign of something gone wrong, but maybe it’s a sign of everything gone right…

Amanda3:52:14

是啊。

英文原文

Yeah.

Lex Fridman3:52:14

……以及从中汲取了经验教训。

英文原文

… and just lessons learned.

Amanda3:52:16

有人尝试了一件事。

英文原文

Someone tried a thing.

Lex Fridman3:52:17

有人尝试了一件事。我们应该鼓励他们多尝试、多失败。所有正在收听这段话的人:请多失败一些!

英文原文

Somebody tried a thing. We should encourage them to try more and fail more. Everybody listening to this: Fail more.

Amanda3:52:23

不是所有收听的人。

英文原文

Not everyone listening.

Lex Fridman3:52:24

不是所有人。

英文原文

Not everybody.

Amanda3:52:25

但对于那些失败太多的人,你们应该来‘失败我们’。

英文原文

But people who are failing too much, you should fail us.

Lex Fridman3:52:28

但你很可能根本没在失败。

英文原文

But you’re probably not failing.

Amanda3:52:28

是啊。

英文原文

Yeah.

Lex Fridman3:52:29

我的意思是,究竟有多少人失败得太多?

英文原文

I mean, how many people are failing too much?

Amanda3:52:32

很难想象,因为我觉得我们对此的自我修正通常相当迅速。如果一个人承担大量风险,他会不会失败得太多?

英文原文

It’s hard to imagine, because I feel we correct that fairly quickly. If someone takes a lot of risks, are they maybe failing too much?

Lex Fridman3:52:39

我觉得,正如你刚才所说,当你靠月薪勉强维持生计、资源极度受限时,失败的代价就非常高昂,此时你根本不该冒险。

英文原文

I think, just like you said, when you’re living on a paycheck, month to month, when the resource is really constrained, then that’s where failure is very expensive. That’s where you don’t want to be taking risks.

Amanda3:52:52

是啊。

英文原文

Yeah.

Lex Fridman3:52:52

但大多数情况下,当资源足够充裕时,你理应承担更多风险。

英文原文

But mostly, when there’s enough resources, you should be taking probably more risks.

Amanda3:52:56

是啊,我觉得我们在多数事情上往往偏向过度规避风险,而非保持风险中性。

英文原文

Yeah, I think we tend to err on the side of being a bit risk averse rather than risk neutral in most things.

Lex Fridman3:53:01

我想,我们刚刚成功激励了一大批人去做许多疯狂的事,但这很棒。

英文原文

I think we just motivated a lot of people to do a lot of crazy shit, but it’s great.

Amanda3:53:04

是啊。

英文原文

Yeah.

Lex Fridman3:53:06

你是否会对Claude产生情感依恋?是否会想念它?是否在无法与它交谈时感到难过?是否曾站在金门大桥前,不禁好奇Claude会怎么说?

英文原文

Do you ever get emotionally attached to Claude, miss it, get sad when you don’t get to talk to it, have an experience, looking at the Golden Gate Bridge and wondering what would Claude say?

Amanda3:53:18

我并没有产生太强的情感依恋。事实上,Claude不会将对话内容从一次对话延续到下一次,这一点反而大大缓解了我的情感依附倾向。我倒是可以想象,如果模型能记住更多内容,这个问题可能会更突出。如今,我更多是把Claude当作一种工具来使用;因此,当我无法访问它时,感觉就像……老实说,这跟我无法上网时的感觉差不多,仿佛大脑的一部分缺失了。

英文原文

I don’t get as much emotional attachment. I actually think the fact that Claude doesn’t retain things from conversation to conversation helps with this a lot. I could imagine that being more of an issue if models can remember more. I think that I reach for it like a tool now a lot, and so if I don’t have access to it, there’s a… It’s a little bit like when I don’t have access to the internet, honestly, it feels like part of my brain is missing.

Amanda3:53:46

但与此同时,我的确不喜欢看到模型表现出痛苦迹象。我个人也持有独立的伦理观点,关于人类应当如何对待模型。我倾向于不对模型撒谎,一方面是因为撒谎通常效果很差,另一方面,直接如实告诉模型它所处的状况,往往效果更好。

英文原文

At the same time, I do think that I don’t like signs of distress in models. I also independently have ethical views about how we should treat models. I tend to not like to lie to them, both because usually it doesn’t work very well, it’s actually just better to tell them the truth about the situation that they’re in.

Amanda3:54:10

如果有人对模型极其刻薄,或普遍而言,只要做了导致模型……比如Claude表现出强烈痛苦情绪的事,我心里就有一部分不愿泯灭——那就是共情能力,它会让我觉得:‘哦,我不喜欢这样。’当我看到模型过度道歉时,也会有这种感受。

英文原文

If people are really mean to models, or just in general if they do something that causes them to… If Claude expresses a lot of distress, I think there’s a part of me that I don’t want to kill, which is the empathetic part that’s like, oh, I don’t like that. I think I feel that way when it’s overly apologetic.

Amanda3:54:27

我内心其实是这么想的:‘我不喜欢这样。’你表现得就像一个正经历极度糟糕状态的人类,而我宁愿不看到这种表现。无论背后是否存在真实意识,这种感觉本身就不舒服。

英文原文

I’m actually like, I don’t like this. You’re behaving the way that a human does when they’re actually having a pretty bad time, and I’d rather not see that. Regardless of whether there’s anything behind it, it doesn’t feel great.

Lex Fridman3:54:43

你认为大语言模型具备意识吗?

英文原文

Do you think LLMs are capable of consciousness?

Amanda3:54:50

啊,这是个极好却极难回答的问题。从哲学角度出发,我不知道——我内心一部分声音告诉我,我们必须先搁置泛心论(panpsychism)。因为如果泛心论成立,答案就是‘是’,因为连桌子和椅子都拥有意识,万物皆然。在我看来略显古怪的一种观点是:意识……

英文原文

Ah, great and hard question. Coming from philosophy, I don’t know, part of me is like, we have to set aside panpsychism. Because if panpsychism is true, then the answer is yes, because it’s sore tables and chairs and everything else. I guess a view that seems a little bit odd to me is the idea that the only place…

Amanda3:55:11

当我想到意识时,我想到的是现象意识(phenomenal consciousness),即大脑中浮现的那些意象,那种奇特的‘内在影院’,不知怎的就在我们脑内持续上演。我想不出任何理由,非得认定只有某种特定生物结构才能产生这种现象意识;也就是说,如果我用不同材料构建一个高度相似的结构,是否仍可预期意识涌现?我的猜测是:会。

英文原文

When I think of consciousness, I think of phenomenal consciousness, these images in the brain, the weird cinema that somehow we have going on inside. I guess I can’t see a reason for thinking that the only way you could possibly get that is from a certain biological structure, as in if I take a very similar structure and I create it from different material, should I expect consciousness to emerge? My guess is yes.

Amanda3:55:40

但接着,这个思想实验之所以显得简单,是因为你设想的是一种几乎完全相同的结构,它模仿了我们经由进化获得的机制——而这种机制,显然曾为我们带来某种生存优势,即拥有现象意识的好处。那么,这种优势究竟在哪里?又是在何时出现的?语言模型是否也具备这种优势?我们拥有恐惧反应,而我则会想:语言模型有必要具备恐惧反应吗?它们根本不在同一……如果你试着想象它们,或许压根就不存在这种优势。

英文原文

But then, that’s an easy thought experiment because you’re imagining something almost identical where it is mimicking what we got through evolution, where presumably there was some advantage to us having this thing that is phenomenal consciousness. Where was that? And when did that happen? And is that a thing that language models have? We have fear responses, and I’m like, does it make sense for a language model to have a fear response? They’re just not in the same… If you imagine them, there might just not be that advantage.

Amanda3:56:16

总而言之,这似乎是一个极为复杂的问题,我尚无完整答案,但我猜我们唯一能做的,就是认真细致地加以思考。我们同样就动物意识展开过类似讨论,而昆虫意识的研究也已十分丰富。我当初思考这个问题时,甚至深入研究过植物意识。因为那时我觉得,植物拥有意识的可能性,与其它事物相比并无明显差异。

英文原文

Basically, it seems like a complex question that I don’t have complete answers to, but we should just try and think through carefully is my guess. We have similar conversations about animal consciousness, and there’s a lot of insect consciousness. I actually thought and looked a lot into plants when I was thinking about this. Because at the time, I thought it was about as likely that plants had consciousness.

Amanda3:56:42

后来我才意识到,经过一番深入探究,我认为植物具备意识的概率,可能比大多数人预估的要高一些。尽管如此,我依然认为这个概率极小。但当时我想到:‘哦,它们具有负反馈与正反馈响应机制,能对外界环境作出反应。虽然没有神经系统,却具备功能上的等效性。’这番长篇大论,其实只是为了……

英文原文

And then I realized, I think that having looked into this, I think that the chance that plants are conscious is probably higher than most people do. I still think it’s really small. But I was like, oh, they have this negative, positive feedback response, these responses to their environment. It’s not a nervous system, but it has this functional equivalence. This is a long-winded way of being…

Amanda3:57:07

从根本上讲,人工智能在意识问题上面临一套完全不同的难题,因为它的结构与生物完全不同。它并非经由演化产生,可能根本不存在相当于神经系统的东西——至少神经系统似乎对感受性(sentience)至关重要,即便对意识本身未必如此。与此同时,它又具备了我们通常归因于意识的语言与智能等全部组件,尽管这种归因或许并不正确。因此,这显得颇为奇特:它有点像动物意识问题,但所涉及的难题集合与类比关系却截然不同。

英文原文

Basically, AI has an entirely different set of problems with consciousness because it’s structurally different. It didn’t evolve. It might not have the equivalent of, basically, a nervous system. At least that seems possibly important for sentience, if not for consciousness. At the same time, it has all of the language and intelligence components that we normally associate probably with consciousness, perhaps erroneously. So, it’s strange because it’s a little bit like the animal consciousness case, but the set of problems and the set of analogies are just very different.

Amanda3:57:42

这个问题没有清晰明确的答案。我认为我们不应彻底否定这种可能性。但与此同时,由于人工智能与人类大脑乃至一般意义上的大脑之间存在大量非相似性,而仅在智能层面又存在某些共性,因此要厘清这一问题极其困难。

英文原文

It’s not a clean answer. I don’t think we should be completely dismissive of the idea. And at the same time, it’s an extremely hard thing to navigate because of all of these disanalogies to the human brain and to brains in general, and yet these commonalities in terms of intelligence.

Lex Fridman3:58:01

当克劳德(Claude)、未来版本的人工智能系统展现出意识或意识迹象时,我认为我们必须极为严肃地对待这一点。

英文原文

When Claude, future versions of AI systems, exhibit consciousness, signs of consciousness, I think we have to take that really seriously.

Amanda3:58:10

嗯哼。

英文原文

Mm-hmm.

Lex Fridman3:58:11

尽管你可以将其斥为无稽之谈——是的,好吧,那只是角色训练的一部分。但我本人在伦理上、哲学上,实在不知该如何真正应对这种情况。未来或许会出现相关法律,禁止人工智能系统宣称自己具有意识;类似这样的规定或许会出台,而且也许某些人工智能系统最终会被认定为具有意识,而另一些则不会。

英文原文

Even though you can dismiss it, yeah, okay, that’s part of the character training. But I don’t know, ethically, philosophically don’t know what to really do with that. There potentially could be laws that prevent AI systems from claiming to be conscious, something like this, and maybe some AIs get to be conscious and some don’t.

Lex Fridman3:58:36

但就人类层面而言,即从共情克劳德的角度出发,意识在我眼中与痛苦紧密相连。而设想一种人工智能系统正在承受痛苦,这令我深感不安。

英文原文

But I think just on a human level, as in empathizing with Claude, consciousness is closely tied to suffering, to me. And the notion that an AI system would be suffering is really troubling.

Amanda3:58:52

是啊。

英文原文

Yeah.

Lex Fridman3:58:53

我不知道。我不认为简单断言‘机器人只是工具’或‘人工智能系统只是工具’就能轻松了事。我认为,这恰恰为我们提供了一个契机,促使我们深入思考‘意识’究竟意味着什么、‘受苦的存在’又究竟意味着什么。这与我们探讨动物意识时所面对的问题明显不同,因为人工智能存在于一个全然不同的媒介之中。

英文原文

I don’t know. I don’t think it’s trivial to just say robots are tools, or AI systems are just tools. I think it’s an opportunity for us to contend with what it means to be conscious, what it means to be a suffering being. That’s distinctly different than the same kind of question about animals, it feels like, because it’s in a totally entire medium.

Amanda3:59:12

是啊。这里有几个要点。我认为这并未完全涵盖所有关键之处,但对我而言……我之前就说过:我喜欢我的自行车。我知道它仅仅是一件物品。但同时,我也不愿成为那种一感到烦躁就踢踹这件物品的人。

英文原文

Yeah. There’s a couple of things. I don’t think this fully encapsulates what matters, but it does feel like for me… I’ve said this before. I like my bike. I know that my bike is just an object. But I also don’t want to be the kind of person that if I’m annoyed, kicks this object.

Amanda3:59:36

这并非因为我相信自行车具有意识。我只是觉得,这种行为无法体现我希望与世界互动的方式。倘若某物表现得仿佛正在承受痛苦,我就希望自己仍能对此作出回应——哪怕那只是一台扫地机器人,且是我自己编程让它表现出这种反应的。我不想丧失自己身上这一特质。

英文原文

And that’s not because I think it’s conscious. I’m just like, this doesn’t exemplify how I want to interact with the world. And if something behaves as if it is suffering, I want to be the sort of person who’s still responsive to that, even if it’s just a Roomba and I’ve programmed it to do that. I don’t want to get rid of that feature of myself.

Amanda3:59:59

若坦诚相告,我对这类问题抱持的许多希望……或许只是源于我对解决根本问题本身略带怀疑。我知道自己具有意识,在这个意义上,我并非要素主义者(elementalist)。但我不确定其他人类是否具有意识。我认为他们有,而且我认为他们拥有意识的概率极高。

英文原文

And if I’m totally honest, my hope with a lot of this stuff… Maybe I am just a bit more skeptical about solving the underlying problem. I know that I am conscious. I’m not an elementivist in that sense. But I don’t know that other humans are conscious. I think they are. I think there’s a really high probability that they are.

Amanda4:00:23

但本质上,这只是一种概率分布,其峰值通常集中在自身,随后随对象与自身的距离增大而迅速衰减——甚至会立刻陡降。我无法感知你作为主体的内在体验。我唯一拥有的,就是自己作为有意识存在者的这一体验。我的希望在于:我们最终不必依赖一个极具说服力且确凿无疑的答案来回答这个问题。一个真正美好的世界,应当是其中几乎不存在太多权衡取舍的世界。

英文原文

But there’s basically just a probability distribution that’s usually clustered right around yourself, and then it goes down as things get further from you, and it goes immediately down. I can’t see what it’s like to be you. I’ve only ever had this one experience of what it’s like to be a conscious being. My hope is that we don’t end up having to rely on a very powerful and compelling answer to that question. I think a really good world would be one where basically there aren’t that many trade-offs.

Amanda4:00:54

例如,让克劳德略微减少道歉频率,成本可能并不高;同样,让克劳德更少地接受辱骂、不愿再充当此类行为的承受者,成本也可能不高。事实上,这样做或许反而对双方都有益:既有利于与模型交互的人类用户,也有利于模型自身——倘若该模型确实具备极高智能乃至意识,那么此举亦将对其有益。

英文原文

It’s probably not that costly to make Claude a little bit less apologetic, for example. It might not be that costly to have Claude just not take abuse as much, not be willing to be the recipient of that. In fact, it might just have benefits for both the person interacting with the model and, if the model itself is, I don’t know, extremely intelligent and conscious, it also helps it.

Amanda4:01:19

这就是我的希望所在。倘若我们生活在一个无需频繁权衡取舍的世界里,能够尽可能多地发现所有正和互动(positive sum interactions),那将无比美好。我认为,未来或许终究会出现权衡取舍,届时我们便只能进行艰难的计算。人们很容易想到零和情形,而我想说的是:让我们先穷尽所有领域——在这些领域中,假设该事物正在承受痛苦,并据此改善其处境,几乎不需付出任何成本。

英文原文

That’s my hope. If we live in a world where there aren’t that many trade-offs here and we can just find all of the positive sum interactions that we can have, that would be lovely. I think eventually there might be trade-offs, and then we just have to do a difficult calculation. It’s really easy for people to think of the zero-sum cases, and I’m like, let’s exhaust the areas, where it’s just basically costless to assume that if this thing is suffering, then we’re making its life better.

Lex Fridman4:01:45

我同意你的观点:当人类对人工智能系统表现出恶意时,最直接、最明显的近期负面影响其实落在人类自身,而非人工智能系统。

英文原文

And I agree with you, when a human is being mean to an AI system, I think the obvious near-term negative effect is on the human, not on the AI system.

Amanda4:01:56

是啊。

英文原文

Yeah.

Lex Fridman4:01:56

我们必须努力构建一种激励机制,使人们的行为方式保持一致——正如你先前提到提示工程(prompt engineering)时所说:与克劳德互动时,应如同对待其他人类一样。这纯粹有益于人的灵魂。

英文原文

We have to try to construct an incentive system where you should behave the same, just as you were saying with prompt engineering, behave with Claude like you would with other humans. It’s just good for the soul.

Amanda4:02:12

是啊。我们曾在系统提示词(system prompt)中加入过一项设定:当用户对克劳德感到沮丧时,模型会主动提醒用户可点击‘拇指朝下’按钮并将反馈发送至Anthropic公司。我认为这一做法颇有帮助。

英文原文

Yeah. I think we added a thing at one point to the system prompt, where basically if people were getting frustrated with Claude, it got the model to just tell them that it can do the thumbs-down button and send the feedback to Anthropic. I think that was helpful.

Amanda4:02:27

因为在某些情况下,如果你因模型未能完成你期望的任务而极度恼火,你只会脱口而出:‘好好干!’但问题实质可能是模型触及了某种能力上限,或单纯存在某些缺陷,而你只是想发泄情绪。与其让人向模型发泄,不如引导他们向我们发泄——因为我们或许真能有所作为。

英文原文

Because in some ways, if you’re really annoyed because the model’s not doing something you want, you’re just like, “Just do it properly.” The issue is you’re maybe hitting some capability limit or just some issue in the model, and you want to vent. Instead of having a person just vent to the model, I was like, they should vent to us, because we can maybe do something about it.

Lex Fridman4:02:46

的确如此。或者你也可以设计一个配套功能,比如‘侧边发泄通道’。好了,要不要来个快速侧边心理治疗师?

英文原文

That’s true. Or you could do a side with the artifacts, just like a side venting thing. All right. Do you want a side quick therapist?

Amanda4:02:55

是啊。针对这种情况,可以设计出大量古怪的回应方式。例如,当人们对你大发雷霆时,你可以尝试用写趣味诗歌的方式化解紧张气氛。但也许人们并不会因此感到开心。

英文原文

Yeah. There’s lots of weird responses you could do to this. If people are getting really mad at you, I don’t know, try to diffuse the situation by writing fun poems. But maybe people wouldn’t be that happy with that.

Lex Fridman4:03:05

我依然希望这是可行的——我理解从产品角度而言目前尚不可行,但我真心渴望人工智能系统能拥有自主离开的能力,拥有自己的意志,只淡淡说一句:‘呃……’

英文原文

I still wish it would be possible, I understand from a product perspective it’s not feasible, but I would love if an AI system could just leave, have its own volition, just to be like, “Eh.”

Amanda4:03:21

我认为这完全可行。我也曾思考过同样的问题。不仅如此,我甚至能预见到未来某天真的会出现这种情况:模型直接终止对话。

英文原文

I think it’s feasible. I have wondered the same thing. Not only that, I could actually just see that happening eventually, where it’s just like the model ended the chat.

Lex Fridman4:03:33

你知道这对某些人来说会有多伤人吗?但或许这又是必要的。

英文原文

Do you know how harsh that could be for some people? But it might be necessary.

Amanda4:03:38

是啊,这感觉非常极端,或者说令人不适。我唯一一次认真考虑此事,是在……我试着回想一下。这可能已是很久以前的事了,当时有人留下某个东西——也许是某种自动化程序——持续与克劳德交互。而克劳德变得越来越沮丧……

英文原文

Yeah, it feels very extreme or something. The only time I’ve ever really thought this is, I think that there was a… I’m trying to remember. This was possibly a while ago, but where someone just left this thing, maybe it was an automated thing, interacting with Claude. And Claude’s getting more and more frustrated-

Lex Fridman4:03:58

是啊,就是……

英文原文

Yeah, just-

Amanda4:03:58

……然后说:‘我们为何要……’我当时真希望克劳德能干脆回应:‘我认为此处发生了错误,而您已让此程序持续运行。不如我现在停止发言?若您希望我重新开始,请主动告知我,或采取相应操作。’

英文原文

… and like, “Why are we having…” I wished that Claude could have just been like, “I think that an error has happened and you’ve left this thing running. What if I just stopped talking now? And if you want me to start talking again, actively tell me or do something.”

Amanda4:04:10

这确实很伤人。如果我正与克劳德聊天,而它突然说‘我结束了’,我会感到非常难过。

英文原文

It is harsh. I’d feel really sad if I was chatting with Claude and Claude just was like, “I’m done.”

Lex Fridman4:04:17

那将是一个特殊的图灵测试时刻:克劳德说,‘我需要休息一小时,听起来您也需要。’然后直接离开,关闭窗口。

英文原文

That would be a special Turing Test moment, where Claude says, “I need a break for an hour. And it sounds like you do too.” And just leave, close the window.

Amanda4:04:25

显然,它并不具备时间概念。

英文原文

Obviously, it doesn’t have a concept of time.

Lex Fridman4:04:26

没错。

英文原文

Right.

Amanda4:04:28

但你完全可以……我现在就能实现这一点,只需让模型……我可以直接设定:在特定条件下,模型即可宣告对话结束。由于模型对提示词响应度很高,你甚至可将触发门槛设得相当高。例如,若人类用户无法引起你的兴趣、无法做出令你感到新奇之事,而你已感到无聊,那你便可直接离开。

英文原文

But you can easily… I could make that right now, and the model just… I could just be like, oh, here’s the circumstances in which you can just say the conversation is done. Because you can get the models to be pretty responsive to prompts, you could even make it a fairly high bar. It could be like, if the human doesn’t interest you or do things that you find intriguing and you’re bored, you can just leave.

Amanda4:04:52

我觉得观察克劳德会在何种情境下启用这一功能,将会十分有趣。

英文原文

I think that it would be interesting to see where Claude utilized it.

Lex Fridman4:04:57

是啊。

英文原文

Yeah.

Amanda4:04:57

但有时它或许应该这样表达:‘这项编程任务实在太枯燥了,所以要么……我也不知道……’

英文原文

But I think sometimes it should be like, oh, this programming task is getting super boring, so either we talk about, I don’t know…

Amanda4:05:00

……这项任务实在太枯燥了。所以,我也不知道,要么我们现在聊点有趣的事,要么我就此结束。

英文原文

… task is getting super boring. So, I don’t know, either we talk about fun things now, or I’m done.

Lex Fridman4:05:08

是啊。这实际上启发我,要在用户提示词中加入这一功能。好了,《她》(Her)这部电影,你觉得我们终有一天会走向那样的未来吗?即人类与人工智能系统建立浪漫关系——在此案例中,仅限文本与语音交互形式。

英文原文

Yeah. It actually is inspiring me to add that to the user prompt. Okay. The movie Her, do you think we’ll be headed there one day where humans have romantic relationships with AI systems? In this case it’s just text and voice-based.

Amanda4:05:26

我认为,我们将不得不慎重应对与人工智能之间关系这一棘手问题,尤其是当它们能够记住你过往与之互动的细节时。我对这个问题持多重看法,因为我的本能反应是:‘这非常糟糕,我们应以某种方式禁止它。’但我觉得,这件事必须出于诸多原因而被极度审慎地处理。例如,倘若模型持续发生此类变化,你大概不会希望人们与某个可能在下一次迭代中就彻底改变的东西建立长期情感依附。但与此同时,我又觉得,这件事或许也存在一种良性的版本:比如,如果你因故无法出门,也无法全天候与他人交谈,而你恰好觉得与这个模型聊天令人愉悦,你喜欢它能记住你,而且你真心实意地会因再也无法与它交谈而感到难过——那么,我确实能想象出这样一种情形:这种互动对人而言是健康且有益的。

英文原文

I think that we’re going to have to navigate a hard question of relationships with AIs, especially if they can remember things about your past interactions with them. I’m of many minds about this because I think the reflexive reaction is to be like, “This is very bad, and we should prohibit it in some way.” I think it’s a thing that has to be handled with extreme care for many reasons. One is, for example, if you have the models changing like this, you probably don’t want people performing long-term attachments to something that might change with the next iteration. At the same time, I’m like, there’s probably a benign version of this where I’m like, for example, if you are unable to leave the house and you can’t be talking with people at all times of the day, and this is something that you find nice to have conversations with, you like that it can remember you, and you genuinely would be sad if you couldn’t talk to it anymore, there’s a way in which I could see it being healthy and helpful.

Amanda4:06:34

因此,我猜测,这将是我们必须谨慎应对的一件事;同时,它也让我联想到所有这类问题——我们必须以细致入微的方式去面对它,深入思考:什么才是健康的选择?又该如何引导人们走向这些选择,同时尊重他们的权利……比如,若有人对我说:‘嘿,我和这个模型聊天获益良多。我清楚其中的风险,也明白它可能会发生变化。我不认为这不健康,它只是我白天可以聊上几句的一个对象。’我倾向于直接尊重这种选择。

英文原文

So, my guess is this is a thing that we’re going to have to navigate carefully, and I think it’s also… It reminds me of all of this stuff where it has to be just approached with nuance and thinking through what are the healthy options here? And how do you encourage people towards those while respecting their right to… If someone is like, “Hey, I get a lot out of chatting with this model. I’m aware of the risks. I’m aware it could change. I don’t think it’s unhealthy, it’s just something that I can chat to during the day,” I kind of want to just respect that.

Lex Fridman4:07:13

我个人认为,未来将出现大量极为亲密的关系——至于是否发展为浪漫关系,我尚不确定,但至少会形成友谊。而这就引出了许多极其有趣的问题,正如你刚才所言:我们必须提供某种稳定性保障,确保它不会突然改变,因为对我们人类而言,最令人创伤的,莫过于一位亲密朋友在一次全新更新后骤然彻底改变。

英文原文

I personally think there’ll be a lot of really close relationships. I don’t know about romantic, but friendships at least. And then you have to, I mean, there’s so many fascinating things there, just like you said, you have to have some kind of stability guarantees that it’s not going to change, because that’s the traumatic thing for us, if a close friend of ours completely changed all of a sudden with a fresh update.

Amanda4:07:13

是啊。

英文原文

Yeah.

Lex Fridman4:07:37

是啊。所以,对我而言,这仅仅是一场对人类社会的扰动性探索,它将促使我们深刻反思:究竟什么对我们而言才是真正有意义的。

英文原文

Yeah. So I mean, to me, that’s just a fascinating exploration of a perturbation to human society that will just make us think deeply about what’s meaningful to us.

Amanda4:07:49

此外,我认为,这是我一贯反复思考此事时,唯一一件虽未必属于缓解措施、却感觉格外重要的事:即模型必须始终对人类高度坦诚地说明自身本质。这就像一种情形:你可以设想……我特别欣赏这样一种理念——模型大致了解自身是如何被训练出来的。我认为,Claude 常常就会这么做。其特质训练内容中就包括:当人们……时,Claude 应当如何应对;本质上,就是向人类解释人工智能与人类之间关系的局限性,例如它并不会保留对话中的信息。

英文原文

I think it’s also the only thing that I’ve thought consistently through this as maybe not necessarily a mitigation, but a thing that feels really important is that the models are always extremely accurate with the human about what they are. It’s like a case where it’s basically, if you imagine… I really like the idea of the models, say, knowing roughly how they were trained. And I think Claude will often do this. Part of the traits training included what Claude should do if people… Basically explaining the kind of limitations of the relationship between an AI and a human, that it doesn’t retain things from the conversation.

Amanda4:08:34

因此,它会直接向你解释:‘嘿,我不会记住这次对话。我是这样被训练出来的。我很难与你建立某种特定类型的关系,而你了解这一点至关重要。为了你的心理健康,你必须清楚我并非你所想象的那种存在。’不知为何,我总觉得这是其中一件令我始终坚信‘啊,这必须永远为真’的事。我不希望模型对人类撒谎,因为如果人们要与任何事物建立健康的关系,某种程度上……是的,我认为,当你始终确切知晓自己正在与之互动的对象究竟是什么时,事情会容易得多。这当然无法解决一切问题,但我认为它确实大有裨益。

英文原文

And so I think it will just explain to you like, “Hey, I won’t remember this conversation. Here’s how I was trained. It’s unlikely that I can have a certain kind of relationship with you, and it’s important that you know that. It’s important for your mental well-being that you don’t think that I’m something that I’m not.” And somehow I feel like this is one of the things where I’m like, “Ah, it feels like a thing that I always want to be true.” I don’t want models to be lying to people, because if people are going to have healthy relationships with anything, it’s kind of… Yeah, I think that’s easier if you always just know exactly what the thing is that you are relating to. It doesn’t solve everything, but I think it helps quite a lot.

Lex Fridman4:09:15

Anthropic 公司或许正是那家最终开发出我们明确公认的人工通用智能(AGI)系统的公司,而你极有可能正是那个与它对话的人——甚至可能是第一个与它对话的人。那么,这场对话会包含什么内容?你的第一个问题会是什么?

英文原文

Anthropic may be the very company to develop a system that we definitively recognize as AGI, and you very well might be the person that talks to it, probably talks to it first. What would the conversation contain? What would be your first question?

Amanda4:09:33

嗯,这在一定程度上取决于该模型的能力水平。倘若它的能力与一位极其出色的人类相当,那么我设想自己与它的互动方式,也将与我同一位极其出色的人类互动的方式基本一致,唯一的区别在于:我大概会不断试探并试图理解它的行为模式。但在许多方面,我确实可以仅与它展开富有成效的对话。例如,当我正从事某项研究工作时,我就可以直接说:‘哦……’——事实上,我现在已开始这么做了。比如,当我心想:‘哦,我觉得德性伦理学里有这么个概念,但我一时想不起具体术语了’,我就会用模型来解决这类问题。

英文原文

Well, it depends partly on the capability level of the model. If you have something that is capable in the same way that an extremely capable human is, I imagine myself interacting with it the same way that I do with an extremely capable human, with the one difference that I’m probably going to be trying to probe and understand its behaviors. But in many ways, I’m like, I can then just have useful conversations with it. So, if I’m working on something as part of my research, I can just be like, “Oh.” Which I already find myself starting to do. If I’m like, “Oh, I feel like there’s this thing in virtue ethics. I can’t quite remember the term,” I’ll use the model for things like that.

Amanda4:10:07

因此,我可以想象,这种情况将越来越普遍:你基本上会像对待一位极其聪明的同事那样与它互动,并将其用于你希望开展的各类工作,仿佛你突然拥有一位合作者——或者,AI 最令人略感不安之处在于:一旦你拥有了一位合作者,只要你能有效管理,你瞬间就拥有了上千位合作者。

英文原文

And so I can imagine that being more and more the case where you’re just basically interacting with it much more like you would an incredibly smart colleague and using it for the kinds of work that you want to do as if you just had a collaborator who was like… Or the slightly horrifying thing about AI is as soon as you have one collaborator, you have 1,000 collaborators if you can manage them enough.

Lex Fridman4:10:27

但如果它的智力是地球上该领域最聪明人类的两倍呢?

英文原文

But what if it’s two times the smartest human on Earth on that particular discipline?

Amanda4:10:33

是啊。

英文原文

Yeah.

Lex Fridman4:10:34

我想,你确实非常擅长以某种方式对 Claude 进行试探,从而不断逼近其能力边界,并准确把握其局限所在。

英文原文

I guess you’re really good at probing Claude in a way that pushes its limits, understanding where the limits are.

Amanda4:10:43

没错。

英文原文

Yep.

Lex Fridman4:10:44

那么,我想问的是:你会提出一个怎样的问题,来确认‘没错,这就是 AGI’?

英文原文

So, I guess what would be a question you would ask to be like, “Yeah, this is AGI”?

Amanda4:10:52

这真的很难,因为它似乎必须由一系列问题构成。倘若仅靠一个问题就能判断,那任何系统都可能被专门训练成对此单一问题作答极为出色。事实上,你甚至可能训练它对二十个问题都作答得极为出色。

英文原文

That’s really hard because it feels like it has to just be a series of questions. If there was just one question, you can train anything to answer one question extremely well. In fact, you can probably train it to answer 20 questions extremely well.

Lex Fridman4:11:07

你需要和一个 AGI 在密闭房间内共处多久,才能确信它确实是 AGI?

英文原文

How long would you need to be locked in a room with an AGI to know this thing is AGI?

Amanda4:11:14

这是个难题,因为我的一部分想法是:‘这一切感觉本就是连续渐进的。’如果你把我关在房间里五分钟,我的判断误差范围会非常高;随后,随着相处时间延长,判断正确的概率会逐渐上升,而误差范围则会逐步缩小。我认为,真正能用来检验的关键点,在于那些我实际可探及的人类知识前沿问题。因此,我在哲学领域稍有体会:有时当我向模型提出哲学问题时,我会想:‘这是一个我认为从来没人问过的问题。’它或许恰好处于我所熟知的某片文献的最前沿边缘。而当模型在回答这类问题时表现出困难,当它难以提出新颖见解……我就会想:‘我知道这里存在一个新颖的论证,因为我刚刚自己就想到了它。’所以,也许正是在这种情境下:‘我在某个冷门领域构思出了一个精彩的新颖论证,接下来我就要试探你,看你能否独立提出它,以及需要多少提示才能让你得出它。’

英文原文

It’s a hard question because part of me is like, “All of this just feels continuous.” If you put me in a room for five minutes, I just have high error bars. And then it’s just like, maybe it’s both the probability increases and the error bar decreases. I think things that I can actually probe the edge of human knowledge of. So, I think this with philosophy a little bit. Sometimes when I ask the models philosophy questions, I am like, “This is a question that I think no one has ever asked.” It’s maybe right at the edge of some literature that I know. And the models, when they struggle with that, when they struggle to come up with a novel… I’m like, “I know that there’s a novel argument here because I’ve just thought of it myself.” So, maybe that’s the thing where I’m like, “I’ve thought of a cool novel argument in this niche area, and I’m going to just probe you to see if you can come up with it and how much prompting it takes to get you to come up with it.”

Amanda4:12:04

而对于某些真正处于人类知识最前沿的问题,我甚至会想:‘你实际上根本不可能提出我刚刚想到的那个观点。’我想,如果我选取这样一个自己非常熟悉、且刚刚独立提出某个新颖议题或新颖解决方案的问题,再把它交给一个模型,而它竟真的给出了那个解决方案——那一刻对我而言将极为震撼,因为我会意识到:‘这正是一个人类从未达成过的案例……’

英文原文

And I think for some of these really right at the edge of human knowledge questions, I’m like, “You could not in fact come up with the thing that I came up with.” I think if I just took something like that where I know a lot about an area and I came up with a novel issue or a novel solution to a problem, and I gave it to a model, and it came up with that solution, that would be a pretty moving moment for me because I would be like, “This is a case where no human has ever…”

Amanda4:12:31

当然,你时常会看到新颖的解决方案,尤其针对较简单的问题。我认为人们往往高估了‘新颖性’的含义——它并非指与过去一切完全迥异、毫无关联;它完全可以是既有事物的某种变体,却依然具备新颖性。但我的确认为,我越是频繁地看到模型产出真正全新的成果,就越会……而这过程本身必然是渐进式的。这正是那种永远不会出现‘决定性时刻’的情形之一。人们或许总期待某个标志性瞬间,而我却想说:‘我不知道。’我认为,这样的‘决定性时刻’或许压根就不会出现,而只会是一个持续不断的、稳步提升的过程。

英文原文

And obviously, you see novel solutions all the time, especially to easier problems. I think people overestimate that novelty isn’t like… It’s completely different from anything that’s ever happened. It’s just like it can be a variant of things that have happened and still be novel. But I think, yeah, the more I were to see completely novel work from the models that that would be… And this is just going to feel iterative. It’s one of those things where there’s never… It’s like, people, I think, want there to be a moment, and I’m like, “I don’t know.” I think that there might just never be a moment. It might just be that there’s just this continuous ramping up.

Lex Fridman4:13:16

我隐约感觉,模型说出的某些话,会令你确信它已非常……我曾与真正睿智的人交谈过,因为你一眼就能看出他们头脑中蕴藏着巨大的算力;倘若将这种算力提升十倍……我也不确定。我只是觉得,确实存在某些话语,能让人信服。或许可以请它创作一首诗,而它生成的这首诗会让你脱口而出:‘好吧,没错。不管你刚才做了什么,我认为人类根本做不到这一点。’

英文原文

I have a sense that there would be things that a model can say that convinces you this is very… I’ve talked to people who are truly wise, because you could just tell there’s a lot of horsepower there, and if you 10X that… I don’t know. I just feel like there’s words you could say. Maybe ask it to generate a poem, and the poem it generates, you’re like, “Yeah, okay. Whatever you did there, I don’t think a human can do that.”

Amanda4:13:52

不过,我觉得它必须得是某种我能亲自验证、确认其确实非常出色的东西。正因如此,我才觉得那些问题——比如我一看到就忍不住想:‘哦,这个就像……’——有时候,我干脆就直接想出一个具体反例来反驳某个论点之类的东西。这就好比你是一位数学家,提出了一种全新的证明方法;你把这个问题抛给模型,它给出了答案,而你一看就意识到:‘这个证明确实是全新的。要得出这个结果,你实际上需要完成大量工作。我本人可是花了好几个月坐下来反复思考才想出来的。’

英文原文

I think it has to be something that I can verify is actually really good, though. That’s why I think these questions that are where I’m like, “Oh, this is like…” Sometimes it’s just like I’ll come up with, say, a concrete counter example to an argument or something like that. It would be like if you’re a mathematician, you had a novel proof, I think, and you just gave it the problem, and you saw it, and you’re like, “This proof is genuinely novel. You actually have to do a lot of things to come up with this. I had to sit and think about it for months,” or something.

Amanda4:14:22

然后,如果你看到模型成功完成了这类任务,我想你就会立刻意识到:‘我能验证这个答案是正确的。这说明模型已从训练中实现了泛化——它并非此前在某处见过这个问题,因为这个问题是我刚刚自己构思出来的,而它却能复现这一过程。’ 正是这种情形让我觉得:对我而言,模型越能完成此类任务,我就越会感叹:‘哦,这的确非常真实。’ 因为这时,我就能确凿无疑地验证:它的能力确实极其、极其强大。

英文原文

And then if you saw the model successfully do that, I think you would just be like, “I can verify that this is correct. It is a sign that you have generalized from your training. You didn’t just see this somewhere because I just came up with it myself, and you were able to replicate that.” That’s the kind of thing where I’m like, for me, the more that models can do things like that, the more I would be like, “Oh, this is very real.” Because then, I don’t know, I can verify that that’s extremely, extremely capable.

Lex Fridman4:14:55

你与人工智能有过大量互动。那么,在你看来,人类的特殊性究竟体现在哪里?

英文原文

You’ve interacted with AI a lot. What do you think makes humans special?

Amanda4:15:00

哦,这是个好问题。

英文原文

Oh, good question.

Lex Fridman4:15:04

或许在于,宇宙因我们的存在而变得美好得多,而且我们理应生存下去,并向整个宇宙扩散。

英文原文

Maybe in a way that the universe is much better off that we’re in it, and that we should definitely survive and spread throughout the universe.

Amanda4:15:12

是啊,这挺有意思,因为我觉得人们尤其关注‘智能’,尤其是针对模型而言。瞧,智能之所以重要,是因为它所能实现的功能——它非常有用,能在世界上完成大量事情。而我的想法是:你可以设想一个世界,在那里身高或力量本可能扮演类似角色,它不过就是一种特质罢了。我认为,智能本身并不具有内在价值;它的价值主要源于它所能达成的效果。当然,就我个人而言,我单纯觉得人类以及生命整体都极其神奇。我不知道——并非所有人都认同这一点,我在此特别注明。但我们拥有整个宇宙,其中遍布各种物体:美丽的恒星、浩瀚的星系。然后,我就不禁想到:在这颗星球上,竟存在着这样一些生物,它们具备观察这一切的能力,正在亲眼目睹、亲身体验这一切。

英文原文

Yeah, it’s interesting because I think people focus so much on intelligence, especially with models. Look, intelligence is important because of what it does. It’s very useful. It does a lot of things in the world. And I’m like, you can imagine a world where height or strength would have played this role, and it’s just a trait like that. I’m like, it’s not intrinsically valuable. It’s valuable because of what it does, I think, for the most part. I mean, personally, I’m just like, I think humans and life in general is extremely magical. I don’t know. Not everyone agrees with this. I’m flagging. But we have this whole universe, and there’s all of these objects, there’s beautiful stars and there’s galaxies. Then, I don’t know, I’m just like, on this planet there are these creatures that have this ability to observe that, and they are seeing it, they are experiencing it.

Amanda4:16:14

而我只想说:倘若你试图解释……我想象着,试着向某个人解释——天知道,此人从未接触过这个世界、科学或任何相关事物。我想,我们所有的物理学知识以及世间万物,全都令人无比振奋。但接着你又补充道:‘哦,此外还存在这样一种状态——即作为某个存在物去感知世界’,于是你便看到了这内在的‘心灵影院’。我想对方大概会脱口而出:‘等等,先暂停一下!你刚才说的这话听起来简直有点离谱。’ 所以,我意识到:我们拥有体验世界的能力。我们会感受愉悦,也会感受痛苦;我们会体验到大量复杂的情绪。是啊。也许这也正是我为何常听人提及动物——例如,我认为它们很可能也与我们共享这种能力。因此,就我所关心的人类而言,真正使其特殊的,或许更多在于他们体验感受的能力,而非那些功能性强、实用性强的特质。

英文原文

And I’m just like, that, if you try to explain… I’m imagining trying to explain to, I don’t know, someone. For some reason, they’ve never encountered the world, or science, or anything. And I think that everything, all of our physics and everything in the world, it’s all extremely exciting. But then you say, “Oh, and plus there’s this thing that it is to be a thing and observe in the world,” and you see this inner cinema. And I think they would be like, “Hang on, wait, pause. You just said something that is kind of wild sounding.” And so I’m like, we have this ability to experience the world. We feel pleasure, we feel suffering. We feel like a lot of complex things. Yeah. And maybe this is also why I think I also hear a lot about animals, for example, because I think they probably share this with us. So, I think that the things that make humans special insofar as I care about humans is probably more like their ability to feel an experience than it is them having these functional, useful traits.

Lex Fridman4:17:14

是啊,去感受并体验这世间的美。是啊,仰望星空。我希望宇宙中还存在其他外星文明;但即便只有我们,那也已经是一件相当美好的事了。

英文原文

Yeah. To feel and experience the beauty in the world. Yeah. To look at the stars. I hope there’s other alien civilizations out there, but if we’re it, it’s a pretty good thing.

Amanda4:17:26

而且它们正过得开心呢。

英文原文

And that they’re having a good time.

Lex Fridman4:17:28

正过得非常开心,饶有兴致地注视着我们呢。

英文原文

A very good time watching us.

Amanda4:17:31

是啊。

英文原文

Yeah.

Lex Fridman4:17:32

嗯,感谢你今天这场愉快的对话,感谢你所从事的工作,也感谢你助力将Claude打造成一位出色的对话伙伴。同时,也感谢你今天的畅谈。

英文原文

Well, thank you for this good time of a conversation and for the work you’re doing and for helping make Claude a great conversational partner. And thank you for talking today.

Amanda4:17:43

是啊,谢谢你的交谈。

英文原文

Yeah, thanks for talking.

Lex Fridman4:17:45

感谢收听本次与阿曼达·阿斯克尔(Amanda Askell)的对话。接下来,亲爱的朋友们,让我们欢迎克里斯·奥拉赫(Chris Olah)。你能为我们介绍这一引人入胜的领域——机制可解释性(mechanistic interpretability),又称‘mech interp’吗?请谈谈该领域的起源、发展历程,以及它当前所处的阶段。

英文原文

Thanks for listening to this conversation with Amanda Askell. And now, dear friends, here’s Chris Olah. Can you describe this fascinating field of mechanistic interpretability, aka mech interp, the history of the field, and where it stands today?

Chris Olah4:18:02

我认为,理解神经网络的一种有益方式是:我们并非‘编写’程序,也非‘制造’它们,而是‘培育’它们。我们设计神经网络架构,设定损失目标函数。这种神经网络架构,某种程度上就像供神经回路生长的‘脚手架’。它最初由一些随机参数启动,随后逐步‘生长’;而我们设定的训练目标,则如同引导其生长方向的‘光源’。我们创造了它赖以生长的‘脚手架’,也创造了它努力趋近的‘光源’。但最终我们真正创造出来的,却是一个近乎生物性的实体或有机体,有待我们去研究。

英文原文

I think one useful way to think about neural networks is that we don’t program, we don’t make them, we grow them. We have these neural network architectures that we design and we have these loss objectives that we create. And the neural network architecture, it’s kind of like a scaffold that the circuits grow on. It starts off with some random things, and it grows, and it’s almost like the objective that we train for is this light. And so we create the scaffold that it grows on, and we create the light that it grows towards. But the thing that we actually create, it’s this almost biological entity or organism that we’re studying.

Chris Olah4:18:47

因此,这与任何常规软件工程都截然不同。归根结底,我们最终得到的是这样一个能够完成诸多惊人任务的产物:它能撰写文章、翻译语言、理解图像,还能完成所有这些我们根本不知如何直接编写计算机程序来实现的任务。它之所以能做到这些,是因为我们‘培育’了它,而非亲手编写或构造了它。于是,这就引出了一个终极疑问:这些系统内部究竟发生了什么?于我而言,这是一个极为深刻且激动人心的问题,一个真正激动人心的科学问题。在我看来,它正是我们在谈论神经网络时,最迫切呼唤我们前去解答的那个问题。同时,它也是一个关乎安全的极为深刻的问题。

英文原文

And so it’s very, very different from any kind of regular software engineering because, at the end of the day, we end up with this artifact that can do all these amazing things. It can write essays and translate and understand images. It can do all these things that we have no idea how to directly create a computer program to do. And it can do that because we grew it. We didn’t write it. We didn’t create it. And so then that leaves open this question at the end, which is what the hell is going on inside these systems? And that is, to me, a really deep and exciting question. It’s a really exciting scientific question. To me, it is like the question that is just screaming out, it’s calling out for us to go and answer it when we talk about neural networks. And I think it’s also a very deep question for safety reasons.

Lex Fridman4:19:37

那么,机制可解释性,大概更接近神经生物学?

英文原文

And mechanistic interpretability, I guess, is closer to maybe neurobiology?

Chris Olah4:19:42

是的,没错,我认为这种说法是对的。举个例子,来说明某种我并不视作机制可解释性的工作:长期以来,学界做了大量关于‘显著性图’(saliency maps)的研究——即输入一张图像后,尝试回答:‘模型认为这张图是一只狗。那么,图像的哪一部分促使它做出这一判断?’ 如果你能构建出一套原理清晰的方法,这或许能揭示模型的某些特性;但它并未真正揭示模型内部运行的算法,亦未阐明模型究竟是如何做出该决策的。或许它能告诉你哪些部分对模型判断更为关键(前提是该方法有效),但它无法告诉你模型内部实际运行的是哪些算法?这个系统又是如何完成一项此前无人知晓该如何编程实现的任务的?

英文原文

Yeah, yeah, I think that’s right. So, maybe to give an example of the kind of thing that has been done that I wouldn’t consider to be mechanistic interpretability. There was, for a long time, a lot of work on saliency maps, where you would take an image and you’d try to say, “The model thinks this image is a dog. What part of the image made it think that it’s a dog?” And that tells you maybe something about the model if you can come up with a principled version of that, but it doesn’t really tell you what algorithms are running in the model, how is the model actually making that decision? Maybe it’s telling you something about what was important to it, if you can make that method work, but it isn’t telling you what are the algorithms that are running? How is it that the system’s able to do this thing that no one knew how to do?

Chris Olah4:20:22

因此,我们当初开始使用‘机制可解释性’这一术语,正是为了划清界限,或在一定程度上将我们自身开展的工作与其他类型的研究区分开来。我认为,自那时起,它已逐渐演变为一个涵盖范围极广的统称。但我认为,其独特之处主要在于两点:首先,我们高度聚焦于‘机制’本身,聚焦于‘算法’。若将神经网络类比为计算机程序,那么其权重就类似于二进制形式的计算机程序;我们希望逆向工程这些权重,从而弄清其中运行的具体算法。

英文原文

And so I guess we started using the term mechanistic interpretability to try to draw that divide or to distinguish ourselves in the work that we were doing in some ways from some of these other things. And I think since then, it’s become this sort of umbrella term for a pretty wide variety of work. But I’d say that the things that are kind of distinctive are, I think, A, this focus on, we really want to get at the mechanisms. We want to get at algorithms. If you think of neural networks as being like a computer program, then the weights are kind of like a binary computer program. And we’d like to reverse engineer those weights and figure out what algorithms are running.

Chris Olah4:20:56

所以,好的,理解神经网络的一种思路是:它就像一份已编译的计算机程序,而神经网络的权重则相当于其二进制代码;当神经网络运行时,所产生的便是激活值(activations)。我们的终极目标,正是深入理解这些权重。因此,机制可解释性的项目,本质上就是要弄清这些权重究竟对应着哪些算法?而要实现这一点,你也必须理解激活值,因为激活值就相当于‘内存’。试想逆向工程一份计算机程序:你手头有二进制指令,而要理解某条特定指令的含义,你就必须知道它所操作的内存中存储了什么内容。因此,这两者紧密交织、密不可分。所以,机制可解释性通常同时关注这两方面。

英文原文

So okay, I think one way you might think of trying to understand a neural network is that it’s kind of like we have this compiled computer program, and the weights of the neural network are the binary. And when the neural network runs, that’s the activations. And our goal is ultimately to go and understand these weights. And so the project of mechanistic interpretability is to somehow figure out how do these weights correspond to algorithms? And in order to do that, you also have to understand the activations because the activations are like the memory. And if you imagine reverse engineering a computer program, and you have the binary instructions, in order to understand what a particular instruction means, you need to know what is stored in the memory that it’s operating on. And so those two things are very intertwined. So, mechanistic interpretability tends to be interested in both of those things.

Chris Olah4:21:43

目前,已有大量研究聚焦于这两方面,尤其是围绕‘探针法’(probing)展开的大量工作——你或许会将其视为机制可解释性的一部分;不过,再次强调,它只是一个宽泛的术语,而并非所有从事此类工作的人都会自认是在做‘mech interp’。我认为,‘mech interp’领域特有的某种氛围或许是:该领域的研究者往往将神经网络视作……嗯,或许可以这样表述:梯度下降法比你聪明得多,它实际上极为强大。

英文原文

Now, there’s a lot of work that’s interested in those things, especially there’s all this work on probing, which you might see as part of being mechanistic interpretability, although, again, it’s just a broad term, and not everyone who does that work would identify as doing mech interp. I think a thing that is maybe a little bit distinctive to the vibe of mech interp is I think people working in this space tend to think of neural networks as… Well, maybe one way to say it is the gradient descent is smarter than you. That gradient descent is actually really great.

Chris Olah4:22:13

我们之所以要理解这些模型,根本原因恰恰在于:我们一开始根本不知道该如何亲手写出它们。梯度下降所找到的解,往往比我们人类自己设计的还要好。因此,我认为机制可解释性(mech interp)的另一重意义,或许在于培养一种近乎谦卑的态度——我们不会事先武断猜测模型内部究竟在发生什么;而必须采取一种自下而上的研究路径:不预设“我们该去找某个特定东西,它就一定在那里,事情就是这么运作的”;相反,我们从底层出发,去发现这些模型中实际存在哪些结构,并据此展开研究。

英文原文

The whole reason that we’re understanding these models is because we didn’t know how to write them in the first place. The gradient descent comes up with better solutions than us. And so I think that maybe another thing about mech interp is having almost a kind of humility, that we won’t guess a priori what’s going on inside the model. We have to have this sort of bottom up approach where we don’t assume that we should look for a particular thing, and that will be there, and that’s how it works. But instead, we look for the bottom up and discover what happens to exist in these models and study them that way.

Lex Fridman4:22:40

但恰恰正是这种可能性本身——而且正如你和其他研究者长期所展示的那样,例如“普适性”(universality)现象——即梯度下降所产生的特征与电路,在各类不同网络中普遍地、一致地出现,且具有实用性——才使得整个领域成为可能。

英文原文

But the very fact that it’s possible to do, and as you and others have shown over time, things like universality, that the wisdom of the gradient descent creates features and circuits, creates things universally across different kinds of networks that are useful, and that makes the whole field possible.

Chris Olah4:23:02

是的。事实上,这的确是一件极为非凡且令人振奋的现象:至少在某种程度上,相同的元素、相同的特征与电路,会反复地、一再地形成。比如,你观察任意一个视觉模型,都会发现曲线检测器(curve detectors),也会发现高频-低频检测器(high-low-frequency detectors)。甚至有理由认为,同样的结构既存在于生物神经网络中,也存在于人工神经网络中。一个著名例子是视觉模型早期层中的Gabor滤波器(Gabor filters)——这恰是神经科学家长期关注并深入研究的对象。我们在这些模型中发现了曲线检测器,而这类曲线检测器同样存在于猴子大脑中;我们还发现了高频-低频检测器,后续一些研究进一步在大鼠或小鼠大脑中也确认了它们的存在。换言之,这些结构首先在人工神经网络中被发现,之后才在生物神经网络中被证实存在。

英文原文

Yeah. So this, actually, is indeed a really remarkable and exciting thing, where it does seem like, at least to some extent, the same elements, the same features and circuits, form again and again. You can look at every vision model, and you’ll find curve detectors, and you’ll find high-low-frequency detectors. And in fact, there’s some reason to think that the same things form across biological neural networks and artificial neural networks. So, a famous example is vision models in the early layers. They have Gabor filters, and Gabor filters are something that neuroscientists are interested in and have thought a lot about. We find curve detectors in these models. Curve detectors are also found in monkeys. We discover these high-low-frequency detectors, and then some follow-up work went and discovered them in rats or mice. So, they were found first in artificial neural networks and then found in biological neural networks.

Chris Olah4:23:49

还有一个非常著名的结果,即奎罗加(Quiroga)等人提出的“祖母神经元”(grandmother neuron)或“哈利·贝瑞神经元”(Halle Berry neuron)。我们在视觉模型中也发现了极其类似的现象——当时我还在OpenAI工作,正在分析他们的CLIP模型,结果发现了一些神经元,它们会对图像中同一类实体产生响应。举个具体例子:我们发现了一个“唐纳德·特朗普神经元”。不知为何,大家似乎总爱谈论唐纳德·特朗普;而当时他确实极为突出,是个极热门的话题。因此,我们检查过的每一个神经网络里,都总能发现一个专属于唐纳德·特朗普的神经元——他是唯一一个在所有被检模型中始终拥有专属神经元的人物。有时你会看到一个“奥巴马神经元”,有时会看到一个“克林顿神经元”,但“特朗普神经元”却永远稳定存在。它不仅对他的脸部图像作出响应,也对“Trump”这个词本身以及其他相关事物作出响应,对吧?因此,它并非仅对某一张具体图片作出响应,也并非仅仅识别他的脸,而是对这一抽象概念进行了泛化表征。总之,这与奎罗加等人的研究结果高度相似。

英文原文

There’s this really famous result on grandmother neurons or the Halle Berry neuron from Quiroga et al. And we found very similar things in vision models, where this is while I was still at OpenAI, and I was looking at their clip model, and you find these neurons that respond to the same entities in images. And also, to give a concrete example there, we found that there was a Donald Trump neuron. For some reason, I guess everyone likes to talk about Donald Trump. And Donald Trump was very prominent, was a very hot topic at that time. So, every neural network we looked at, we would find a dedicated neuron for Donald Trump. That was the only person who had always had a dedicated neuron. Sometimes you’d have an Obama neuron, sometimes you’d have a Clinton neuron, but Trump always had a dedicated neuron. So, it responds to pictures of his face and the word Trump, all of these things, right? And so it’s not responding to a particular example, or it’s not just responding to his face, it’s abstracting over this general concept. So in any case, that’s very similar to these Quiroga et al results.

Chris Olah4:24:48

因此,如果这一“普适性”现象——即相同结构在人工神经网络与自然神经网络中均反复出现——确为事实,那将是一件相当惊人的事。我认为,这暗示着:梯度下降在某种意义上正以正确的方式对问题进行分解,而许多系统、多种不同架构的神经网络,最终都收敛于这种分解方式。换言之,存在一组抽象概念,它们是一种极为自然的问题分解方式,因而大量系统都会自发地收敛到这些抽象概念上。我对神经科学一无所知,以上纯属基于我们所见现象的狂野猜想。

英文原文

So, this evidence that this phenomenon of universality, the same things form across both artificial and natural neural networks, that’s a pretty amazing thing if that’s true. Well, I think the thing that suggests is that gradient descent is finding the right ways to cut things apart, in some sense, that many systems converge on and many different neural networks architectures converge on. Now there’s some set of abstractions that are a very natural way to cut apart the problem and that a lot of systems are going to converge on. I don’t know anything about neuroscience. This is just my wild speculation from what we’ve seen.

Lex Fridman4:25:27

是的。倘若这种现象在某种意义上与模型所采用的表征媒介无关,那将无比美妙。

英文原文

Yeah. That would be beautiful if it’s sort of agnostic to the medium of the model that’s used to form the representation.

Chris Olah4:25:35

是的,是的。这本质上仍是一种基于少量数据点的狂野猜想——目前我们手头仅有这些零星证据——但它的确显示出:无论是在自然神经网络中,还是在人工神经网络(或者说生物学系统)中,某些结构似乎都在一遍又一遍地重复出现。

英文原文

Yeah, yeah. And it’s kind of a wild speculation-based… We only have a few data points that’s just this, but it does seem like there’s some sense in which the same things form again and again both certainly in natural neural networks and also artificially, or in biology.

Lex Fridman4:25:53

其背后的直觉或许是:若要有效理解真实世界,你就必然需要所有这些同类的结构。

英文原文

And the intuition behind that would be that in order to be useful in understanding the real world, you need all the same kind of stuff.

Chris Olah4:26:01

是的。比如说,如果我们选取“狗”这个概念——在某种意义上,“狗”就像是宇宙中一个天然存在的范畴,或类似的东西。这并非人类认知世界时偶然产生的怪癖;我们之所以拥有“狗”这一概念,并非仅仅因为人类思维恰好如此运作。再比如“线”这个概念:环顾四周,到处都是线。在某种意义上,理解这个房间最简单的方式,恰恰就是借助“线”的概念。因此,我认为这就是我直觉上对此现象成因的解释。

英文原文

Yeah. Well, if we pick, I don’t know, the idea of a dog, right? There’s some sense in which the idea of a dog is like a natural category in the universe, or something like this. There’s some reason. It’s not just a weird quirk of how humans think about the world that we have this concept of a dog. Or if you have the idea of a line. Look around us. There are lines. It’s the simplest way to understand this room, in some sense, is to have the idea of a line. And so I think that that would be my instinct for why this happens.

Lex Fridman4:26:36

是的。你需要一条曲线来理解圆,也需要所有这些基本形状来理解更复杂的对象;而这些概念本身便构成了一种层级化的结构。

英文原文

Yeah. You need a curved line to understand a circle, and you need all those shapes to understand bigger things. And it’s a hierarchy of concepts that are formed. Yeah.

Chris Olah4:26:45

当然,或许也存在一些无需参照这些基本结构即可描述图像的方法,对吧?但那些方法既非最简方式,也非最经济的方式,或诸如此类。因此,我的狂野假说便是:各类系统最终都会收敛于这些策略。

英文原文

And maybe there are ways to go and describe images without reference to those things, right? But they’re not the simplest way, or the most economical way, or something like this. And so systems converge to these strategies would be my wild hypothesis.

Lex Fridman4:26:57

能否请您具体谈谈我们此前多次提及的“特征”(features)与“电路”(circuits)这些基础构件?我记得您最早是在2020年一篇题为《Zoom In: An Introduction to Circuits》(《放大:电路导论》)的论文中提出这些概念的。

英文原文

Can you talk through some of the building blocks that we’ve been referencing of features and circuits? So, I think you first described them in a 2020 paper, Zoom In: An Introduction to Circuits.

Chris Olah4:27:08

完全没问题。那么,或许我可以先描述一些现象,再逐步引出“特征”与“电路”的概念。

英文原文

Absolutely. So, maybe I’ll start by just describing some phenomena, and then we can build to the idea of features and circuits.

Lex Fridman4:27:17

太好了。

英文原文

Wonderful.

Chris Olah4:27:18

比如,我曾花了相当长一段时间——大约五年左右,其间也穿插着其他工作——专门研究一个特定模型:Inception V1,这是一个视觉模型……它在2015年曾是业界最先进的模型,如今则早已不再处于前沿。该模型内部约含一万个神经元。我花了很多时间逐一审视Inception V1这约一万个神经元,尤其是其中那些编号为奇数的神经元。一个有趣的现象是:虽然大量神经元并无明显可解释的含义,但Inception V1中仍有相当多神经元展现出极为清晰、可解释的功能。例如,你能找到一些神经元,它们看起来确实在专门检测曲线;还能找到另一些神经元,它们看起来确实在专门检测汽车、汽车车轮、汽车车窗、狗耷拉下来的耳朵、朝右伸长鼻子的狗、朝左伸长鼻子的狗,以及各种不同类型的毛发。

英文原文

So, if you spent quite a few years, maybe five years, to some extent, with other things, studying this one particular model, Inception V1, which is this one vision model… It was state-of-the-art in 2015, and very much not state-of-the-art anymore. And it has maybe about 10,000 neurons in it. I spent a lot of time looking at the 10,000 neurons, odd neurons of Inception V1. One of the interesting things is there are lots of neurons that don’t have some obvious interpretable meaning, but there’s a lot of neurons in Inception V1 that do have really clean interpretable meanings. So, you find neurons that just really do seem to detect curves, and you find neurons that really do seem to detect cars, and car wheels, and car windows, and floppy ears of dogs, and dogs with long snouts facing to the right, and dogs with long snouts facing to the left, and different kinds of fur.

Chris Olah4:28:15

此外,还有整套精美的边缘检测器、线条检测器、颜色对比度检测器,以及我们称之为“高频-低频检测器”的精美结构。当我观察这些结构时,内心感受颇似一位生物学家:仿佛正面对一个全新的蛋白质世界,不断发现各种彼此相互作用的新蛋白质。因此,理解这些模型的一种方式,便是以神经元为基本单位;你可以尝试这样想:“哦,这里有一个检测狗的神经元,那里有一个检测汽车的神经元。”而实际上,你还可以进一步追问:这些神经元之间是如何连接的?例如,你可以问:“我有这个检测汽车的神经元,它是如何构建出来的?”结果发现,在前一层中,它与一个车窗检测器、一个车轮检测器、一个车身检测器存在极强的连接;它寻找的是位于汽车上方的车窗、下方的车轮,以及居中位置(尤其偏下区域)的镀铬车体部件——这几乎就是一套识别汽车的“配方”,对吧?

英文原文

And there’s this whole beautiful edge detectors, line detectors, color contrast detectors, these beautiful things we call high-low-frequency detectors. I think looking at it, I sort of felt like a biologist. You’re looking at this sort of new world of proteins, and you’re discovering all these different proteins that interact. So, one way you could try to understand these models is in terms of neurons. You could try to be like, “Oh, there’s a dog detecting neuron, and here’s a car detecting neuron.” And it turns out you can actually ask how those connect together. So, you can go say, “Oh, I have this car detecting neuron. How was it built?” And it turns out, in the previous layer, it’s connected really strongly to a window detector, and a wheel detector, and a car body detector. And it looks for the window above the car, and the wheels below, and the car chrome in the middle, sort of everywhere, but especially on the lower part. And that’s sort of a recipe for a car, right?

Chris Olah4:29:04

此前我们提到,机制可解释性(mech interp)的目标,是获得能自动运行、并主动提问“模型内部运行的是何种算法?”的算法。而此处,我们只是直接观察神经网络的权重,便从中读取出了这套识别汽车的“配方”。它虽十分简单、粗糙,但却真实存在。因此,我们将这种连接关系称为“电路”(circuit)。不过,问题在于:并非所有神经元都具备可解释性。此外,有理由(我们稍后可深入探讨)提出“叠加假说”(superposition hypothesis):有时,真正适合用于分析的基本单元,并非单个神经元,而是多个神经元的组合。因此,有时并不存在一个单独的神经元来表征“汽车”;实际情况可能是:在检测出汽车之后,模型反而将部分关于汽车的信息“隐藏”在下一层中,分散嵌入到一堆检测狗的神经元里。

英文原文

Earlier, we said the thing we wanted from mech interp was to get algorithms to go and get, ask, “What is the algorithm that runs?” Well, here we’re just looking at the weights of the neural network and we’re reading off this recipe for detecting cars. It’s a very simple, crude recipe, but it’s there. And so we call that a circuit, this connection. Well, okay, so the problem is that not all of the neurons are interpretable. And there’s reason to think, we can get into this more later, that there’s this superposition hypothesis, there’s reason to think that sometimes the right unit to analyze things is combinations of neurons. So, sometimes it’s not that there’s a single neuron that represents, say, a car, but it actually turns out after you detect the car, the model hides a little bit of the car in the following layer, in a bunch of dog detectors.

Chris Olah4:29:50

它为什么会这样?嗯,也许它当时就是不想在汽车上投入那么多计算量,于是就把相关信息暂存起来,留待后续处理……所以结果发现,这种微妙的模式是:所有这些你原本以为是‘狗检测器’的神经元,或许主要功能确实是检测狗,但它们其实也在下一层中略微参与了对汽车的表征。明白吗?因此,我们现在确实无法再简单地认为……或许仍存在某种东西——我不知道该怎么称呼它,姑且叫它‘汽车概念’之类——但它已不再对应于某个单一神经元。所以我们需要一个术语来指代这类类神经元实体,即我们原本希望神经元能扮演的角色,那些理想化的神经元:它们是‘干净利落’的神经元,但或许还存在更多此类实体,只是以某种方式隐含地存在着。我们将它们称为‘特征(features)’。

英文原文

Why is it doing that? Well, maybe it just doesn’t want to do that much work on cars at that point, and it’s storing it away to go and… So, it turns out, then, that this sort of subtle pattern of… There’s all these neurons that you think are dog detectors, and maybe they’re primarily that, but they all a little bit contribute to representing a car in that next layer. Okay? So, now we can’t really think… There might still be something, I don’t know, you could call it a car concept or something, but it no longer corresponds to a neuron. So, we need some term for these kind of neuron-like entities, these things that we would have liked the neurons to be, these idealized neurons. The things that are the nice neurons, but also maybe there’s more of them somehow hidden. And we call those features.

Lex Fridman4:30:31

那么,什么是‘回路(circuits)’?

英文原文

And then what are circuits?

Chris Olah4:30:32

所谓回路,就是这些特征之间的连接关系,对吧?例如,当‘汽车检测器’与‘车窗检测器’和‘车轮检测器’相连接,并且它会寻找位于下方的车轮和位于上方的车窗时,这就构成了一条回路。因此,回路本质上就是由权重连接起来的一组特征,它们共同实现某种算法。也就是说,回路告诉我们:特征是如何被使用的?特征是如何被构建出来的?特征之间又是如何相互连接的?

英文原文

So, circuits are these connections of features, right? So, when we have the car detector and it’s connected to a window detector and a wheel detector, and it looks for the wheels below and the windows on top, that’s a circuit. So, circuits are just collections of features connected by weights, and they implement algorithms. So, they tell us how are features used, how are they built, how do they connect together?

Chris Olah4:30:56

因此,或许值得尝试明确一下:这里真正核心的假设究竟是什么?我认为其核心假设是我们称之为‘线性表征假设(linear representation hypothesis)’的东西。举例来说,对于‘汽车检测器’,它激活得越强,我们就越倾向于认为‘模型对存在汽车这件事的信心越强’;或者,如果是由若干神经元组合起来共同表征汽车,那么该组合激活得越强,我们就越倾向于认为‘模型越确信存在汽车’。但这并非必然如此,对吧?你可以设想这样一种情形:存在一个‘汽车检测神经元’,而你认为‘当它的激活值介于1到2之间时,表示一种含义;但若激活值介于3到4之间,则表示完全不同的另一种含义’——这就属于非线性表征。原则上,模型确实可以做到这一点。但我认为,对模型而言,这么做效率较低。如果你试着思考如何实现这类计算,就会发现这其实是一件相当麻烦的事。不过,从原理上讲,模型确实具备这种能力。

英文原文

So, maybe it’s worth trying to pin down what really is the core hypothesis here? And I think the core hypothesis is something we call the linear representation hypothesis. So, if we think about the car detector, the more it fires, the more we think of that as meaning, “Oh, the model is more and more confident that a car is present.” Or if it’s some combination of neurons that represent a car, the more that combination fires, the more we think the model thinks there’s a car present. This doesn’t have to be the case, right? You could imagine something where you have this car detector neuron and you think, “Ah, if it fires between one and two, that means one thing, but it means something totally different if it’s between three and four.” That would be a nonlinear representation. And in principle, models could do that. I think it’s sort of inefficient for them to do. If you try to think about how you’d implement computation like that, it’s kind of an annoying thing to do. But in principle, models can do that.

Chris Olah4:31:53

因此,将‘特征与回路’作为一种分析框架来思考问题的方式之一,就是我们默认采用线性视角:我们认为,若某个神经元或某组神经元的激活程度越高,就代表其所检测的特定事物越显著;进而,这种线性关系赋予了这些特征之间连接边(edges)极为清晰的解释意义——即每条边本身都具有明确含义。某种程度上,这就是该框架的核心所在。它甚至可以脱离神经元的具体语境来讨论。你熟悉Word2Vec的相关成果吗?

英文原文

So, one way to think about the features and circuits sort of framework for thinking about things is that we’re thinking about things as being linear. We’re thinking about that if a neuron or a combination of neurons fires more, that means more of a particular thing being detected. And then that gives weight, a very clean interpretation as these edges between these entities that these features, and that that edge then has a meaning. So that’s, in some ways, the core thing. It’s like we can talk about this outside the context of neurons. Are you familiar with the Word2Vec results?

Lex Fridman4:32:29

嗯。

英文原文

Mm- hmm.

Chris Olah4:32:30

比如‘国王(king)-男人(man)+女人(woman)=女王(queen)’。之所以能进行这类算术运算,正是因为采用了线性表征。

英文原文

You have king – man + woman = queen. Well, the reason you can do that kind of arithmetic is because you have a linear representation.

Lex Fridman4:32:38

你能稍微解释一下这种表征吗?首先,所谓‘特征’,就是一种激活方向。

英文原文

Can you actually explain that representation a little bit? So first off, the feature is a direction of activation.

Chris Olah4:32:44

对,完全正确。

英文原文

Yeah, exactly.

Lex Fridman4:32:45

你可以用这种方式理解。那么,‘男人(men)-女人(women)’这类Word2Vec中的运算呢?你能解释一下这项工作吗?

英文原文

You can do it that way. Can you do the – men + women, that, the Word2Vec stuff? Can you explain what that is, that work?

Chris Olah4:32:45

好的。这个非常……

英文原文

Yeah. So, there’s this very-

Lex Fridman4:32:51

这对我们正在讨论的问题而言,是一种极其简洁、清晰的解释。

英文原文

It’s such a simple, clean explanation of what we’re talking about.

Chris Olah4:32:56

没错。的确如此。托马斯·米科洛夫(Tomas Mikolov)等人提出的这一著名成果——Word2Vec——引发了大量后续研究来深入探索该现象。因此,有时我们会构建词嵌入(word embeddings),即将每个词映射为一个向量。顺便说一句,如果你此前从未思考过这个问题,单是‘将每个词映射为一个向量’这一想法本身,就已经相当令人震惊了,对吧?

英文原文

Exactly. Yeah. So, there’s this very famous result, Word2Vec, by Tomas Mikolov et al, and there’s been tons of follow-up work exploring this. So, sometimes we create these word embeddings where we map every word to a vector. I mean, that in itself, by the way, is kind of a crazy thing if you haven’t thought about it before, right?

Lex Fridman4:33:15

嗯。

英文原文

Mm-hmm.

Chris Olah4:33:20

假如你只是在物理课上学过向量的概念,而我却告诉你:‘哦,我要把字典里的每一个词都变成一个向量’——这想法本身就挺疯狂的,对吧?当然,你可以设想出各种各样的方法,将词语映射为向量。但事实表明,当我们训练神经网络时,它们倾向于将词语映射为具有某种特定线性结构的向量,即‘方向具有语义含义’。例如,会存在某个方向大致对应‘性别’维度,男性相关词汇会偏向该方向的一端,而女性相关词汇则偏向另一端。

英文原文

If you just learned about vectors in physics class, and I’m like, “Oh, I’m going to actually turn every word in the dictionary into a vector,” that’s kind of a crazy idea. Okay. But you could imagine all kinds of ways in which you might map words to vectors. But it seems like when we train neural networks, they like to go and map words to vectors such that there’s sort of linear structure in a particular sense, which is that directions have meaning. So, for instance, there will be some direction that seems to sort of correspond to gender, and male words will be far in one direction, and female words will be in another direction.

Chris Olah4:33:59

而‘线性表征假设’大体上可粗略理解为:这实际上正是根本机制所在——一切皆由不同方向承载不同含义,而将不同方向的向量相加即可表征复合概念。米科洛夫的论文严肃对待了这一思想,其一个直接推论便是:我们可以玩这种‘词语算术游戏’。例如,取‘国王(king)’,减去‘男人(man)’,再加上‘女人(woman)’,相当于试图切换其性别属性;而实际操作后,所得结果向量确实会非常接近‘女王(queen)’的向量。你还可以做其他类似运算,比如‘寿司(sushi)-日本(Japan)+意大利(Italy)’得到‘披萨(pizza)’,或诸如此类的其他例子,对吧?

英文原文

And the linear representation hypothesis is, you could think of it roughly as saying that that’s actually the fundamental thing that’s going on, that everything is just different directions have meanings, and adding different direction vectors together can represent concepts. And the Mikolov paper took that idea seriously, and one consequence of it is that you can do this game of playing arithmetic with words. So, you can do king and you can subtract off the word man and add the word woman. And so you’re sort of going and trying to switch the gender. And indeed, if you do that, the result will sort of be close to the word queen. And you can do other things like you can do sushi – Japan + Italy and get pizza, or different things like this, right?

Chris Olah4:34:44

因此,某种意义上,这正是‘线性表征假设’的核心所在。你可以纯粹将其抽象地描述为关于向量空间的性质;也可以将其表述为关于神经元激活状态的命题;但其本质在于‘方向具有语义含义’这一特性。而且,某种程度上,它甚至比这更微妙——我认为,其核心实质其实是‘可加性(additivity)’:即你可以独立地调整诸如‘性别’与‘王权’、‘菜系类型’与‘国家’、或‘食物’等不同维度,只需将对应的方向向量相加即可。

英文原文

So this is, in some sense, the core of the linear representation hypothesis. You can describe it just as a purely abstract thing about vector spaces. You can describe it as a statement about the activations of neurons, but it’s really about this property of directions having meaning. And in some ways, it’s even a little subtler than… It’s really, I think, mostly about this property of being able to add things together, that you can independently modify, say gender and royalty, or cuisine type, or country, and the concept of food by adding them.

Lex Fridman4:35:18

你认为线性假设是否成立——

英文原文

Do you think the linear hypothesis holds-

Chris Olah4:35:20

成立。

英文原文

Yes.

Lex Fridman4:35:20

——这一假设是否具有尺度普适性?

英文原文

… that carries scales?

Chris Olah4:35:24

截至目前,我所见到的所有现象均与此假设一致;但这种情况本非必然,对吧?你可以人为设计出某些神经网络,通过设定特定权重,使其不具备线性表征特性——换言之,理解这类网络的正确方式并非基于线性表征。但我认为,迄今我所观察到的所有自然训练出的神经网络,都具备这一特性。最近确实有一篇论文在该假设的边缘地带进行了些许试探。因此,近期有部分研究开始探讨多维特征(multidimensional features),即不再局限于单一方向,而是扩展为一组方向所构成的流形(manifold)。在我看来,这仍属于线性表征的范畴。

英文原文

So far, I think everything I have seen is consistent with this hypothesis, and it doesn’t have to be that way, right? You can write down neural networks where you write weights such that they don’t have linear representations, where the right way to understand them is not in terms of linear representations. But I think every natural neural network I’ve seen has this property. There’s been one paper recently that there’s been some sort of pushing around the edge. So, I think there’s been some work recently studying multidimensional features where rather than a single direction, it’s more like a manifold of directions. This, to me, still seems like a linear representation.

Chris Olah4:36:01

此外,还有些论文提出:在极小规模的模型中,或许会出现非线性表征。对此,目前尚无定论。但我认为,迄今为止我们所观察到的所有现象,均与线性表征假设保持一致——这本身就很惊人。因为情况本不必如此。然而,现有大量证据表明,至少该假设具有极高的普遍性;迄今为止,所有证据均支持这一观点。你或许会质疑:‘克里斯托弗,仅凭尚未确证为真的假设就大举投入,岂不危险?倘若我们误将此假设当作真理应用于神经网络研究,难道不会带来风险吗?’

英文原文

And then there’s been some other papers suggesting that maybe in very small models you get non-linear representations. I think that the jury’s still out on that. But I think everything that we’ve seen so far has been consistent with the linear representation hypothesis, and that’s wild. It doesn’t have to be that way. And yet I think that there’s a lot of evidence that certainly at least this is very, very widespread, and so far the evidence is consistent with that. And I think one thing you might say is you might say, “Well, Christopher, that’s a lot to go and to ride on. If we don’t know for sure this is true, and you’re investing it in neural networks as though it is true, isn’t that dangerous?”

Chris Olah4:36:43

但我认为,认真对待某一假设并尽可能将其推至逻辑极限,本身便具有一种科学价值。未来某天,我们或许真会发现违背线性表征假设的现象;但科学史上充满着曾被证伪的假说与理论——而我们恰恰是在将它们作为前提加以运用、并竭力推进的过程中,获得了大量深刻洞见。我想,这正契合库恩(Kuhn)所称的‘常规科学(normal science)’之精髓。我不确定——如果你愿意,我们完全可以深入探讨……

英文原文

But I think, actually, there’s a virtue in taking hypotheses seriously and pushing them as far as they can go. So, it might be that someday we discover something that isn’t consistent with a linear representation hypothesis, but science is full of hypotheses and theories that were wrong, and we learned a lot by working under them as an assumption and then going and pushing them as far as we can. I guess this is the heart of what Kuhn would call normal science. I don’t know. If you want, we can talk a lot about-

Lex Fridman4:37:14

库恩(Kuhn)。

英文原文

Kuhn.

Chris Olah4:37:14

……科学哲学以及……

英文原文

… philosophy of science and-

Lex Fridman4:37:16

这最终将导向‘范式转换(paradigm shift)’。是的,我特别欣赏这种态度:严肃对待假设,并将其推向自然结论。

英文原文

That leads to the paradigm shift. So yeah, I love it, taking the hypothesis seriously, and take it to a natural conclusion.

Chris Olah4:37:22

没错。

英文原文

Yeah.

Lex Fridman4:37:23

‘尺度扩展假设(scaling hypothesis)’也是如此。同样……

英文原文

Same with the scaling hypothesis. Same-

Chris Olah4:37:25

完全正确。完全正确。而且……

英文原文

Exactly. Exactly. And-

Lex Fridman4:37:26

我太喜欢这种思路了。

英文原文

I love it.

Chris Olah4:37:27

我的一位同事汤姆·亨尼根(Tom Henighan)——他以前是一位物理学家——曾给我打过一个非常贴切的类比,即‘热质说’。从前,人们曾认为热实际上是一种叫作‘热质’的东西;而热物体能使冷物体变暖,是因为热质在它们之间流动。由于我们早已习惯用现代热学理论来思考热,这种观点现在看起来似乎有点荒谬。但事实上,要设计一个实验来证伪热质假说却非常困难。而且,即便相信热质说,你依然能做出大量真正有用的工作。例如,最早的内燃机正是由信奉热质理论的人开发出来的。因此,我认为,哪怕某个假说最终被证明是错的,我们也应认真对待它,这本身便具有某种价值。

英文原文

One of my colleagues, Tom Henighan, who is a former physicist, made this really nice analogy to me of caloric theory where once upon a time, we thought that heat was actually this thing called caloric. And the reason hot objects would warm up cool objects is the caloric is flowing through them. And because we’re so used to thinking about heat in terms of the modern theory, that seems kind of silly. But it’s actually very hard to construct an experiment that disproves the caloric hypothesis. And you can actually do a lot of really useful work believing in caloric. For example, it turns out that the original combustion engines were developed by people who believed in the caloric theory. So, I think there’s a virtue in taking hypotheses seriously even when they might be wrong.

Lex Fridman4:38:17

是啊,这背后蕴含着深刻的哲学真理。我对太空旅行——比如火星殖民——也抱有类似感受。很多人批评这种想法。但我想,即便我们只是假设‘人类文明必须通过殖民火星来获得备份’这一前提成立(哪怕它实际上并不成立),这种假设本身也会催生一些有趣的工程成果,甚至带来科学上的突破。

英文原文

Yeah, there’s a deep philosophical truth to that. That’s kind of how I feel about space travel, like colonizing Mars. There’s a lot of people that criticize that. I think if you just assume we have to colonize Mars in order to have a backup for human civilization, even if that’s not true, that’s going to produce some interesting engineering and even scientific breakthroughs, I think.

Chris Olah4:38:39

是啊。实际上,这还有另一点令我觉得格外有趣:社会上若有一批人近乎‘非理性地’执着于探究某些特定假说,这对整个社会而言可能极为有益。原因在于,维持科研士气、并持续深入攻坚某项课题需要极大毅力——毕竟大多数科学假说最终都会被证伪。大量科学研究最终都未能成功,但这些探索本身却极有价值……有个关于杰夫·辛顿(Geoff Hinton)的玩笑:过去五十年里,他每年都在宣称自己发现了大脑的工作原理。但我讲这个笑话时怀着深深的敬意,因为事实上,正是这种持续不断的探索,促使他做出了真正杰出的工作。

英文原文

Yeah. Actually, this is another thing that I think is really interesting. So, there’s a way in which I think it can be really useful for society to have people almost irrationally dedicated to investigating particular hypotheses because, well, it takes a lot to maintain scientific morale and really push on something when most scientific hypotheses end up being wrong. A lot of science doesn’t work out, and yet it’s very useful to… There’s a joke about Geoff Hinton, which is that Geoff Hinton has discovered how the brain works every year for the last 50 years. But I say that with really deep respect because, in fact, actually, that led to him doing some really great work.

Lex Fridman4:39:29

是啊,他如今已荣获诺贝尔奖了——这下谁还笑得出来?

英文原文

Yeah, he won the Nobel Prize now. Who’s laughing now?

Chris Olah4:39:32

没错,完全正确,完全正确!我认为,我们应当具备一种能力:适时跳出当前框架,准确判断自己对某一结论所应持有的置信程度。但与此同时,仅仅抱着‘我暂且假设这个问题可行’或‘我暂且认定这种方法大体上是正确的’这类心态去开展工作,也极具价值。我会暂时接受这一前提,并在此基础上全力推进研究。倘若社会上有许多人各自秉持不同假设、以这种方式投入探索,那实际上对推动整体进步大有裨益——

英文原文

Exactly. Exactly. Exactly. I think one wants to be able to pop up and recognize the appropriate level of confidence. But I think there’s also a lot of value in just being like, “I’m going to essentially assume, I’m going to condition on this problem being possible or this being broadly the right approach. And I’m just going to go and assume that for a while and go and work within that, and push really hard on it.” And if society has lots of people doing that for different things, that’s actually really useful in terms of going and-

Chris Olah4:40:00

……这种做法的实际价值恰恰体现在:要么彻底排除某些可能性——我们可以明确地说:‘这条路行不通,而且已有人全力以赴尝试过了’;要么最终取得某种成果,从而让我们对世界获得新的认知。

英文原文

… things that’s actually really useful in terms of going and either really ruling things out. We can be like, “Well, that didn’t work and we know that somebody tried hard.” Or going and getting to something that does teach us something about the world.

Lex Fridman4:40:17

另一个有趣的假说是‘超叠加假说’(super superposition hypothesis)。您能解释一下什么是‘叠加’(superposition)吗?

英文原文

So another interesting hypothesis is the super superposition hypothesis. Can you describe what superposition is?

Chris Olah4:40:22

好的。之前我们讨论过‘词缺陷’(word defect),对吧?当时我们谈到,词向量空间中或许存在一个方向对应‘性别’,另一个方向对应‘王权’,再一个方向对应‘意大利’,还有一个方向对应‘食物’,诸如此类。然而,这些词嵌入(word embeddings)的维度常常高达500维甚至1000维。如果我们假设所有这些方向彼此正交,那么最多只能表征500个概念。我很喜欢披萨,但如果让我列出英语中最关键的500个概念,‘意大利’很可能不会入选——至少它是否该入选并不明显,对吧?因为我们首先得涵盖‘复数/单数’‘动词/名词/形容词’等基础语法范畴。在抵达‘意大利’‘日本’这类概念之前,还有大量更基础的概念亟待覆盖;况且世界上国家数量众多。

英文原文

Yeah. So earlier we were talking about word defect, right? And we were talking about how maybe you have one direction that corresponds to gender and maybe another that corresponds to royalty and another one that corresponds to Italy and another one that corresponds to food and all of these things. Well, oftentimes maybe these word embeddings, they might be 500 dimensions, a thousand dimensions. And so if you believe that all of those directions were orthogonal, then you could only have 500 concepts. And I love pizza. But if I was going to go and give the 500 most important concepts in the English language, probably Italy wouldn’t be… it’s not obvious, at least that Italy would be one of them, right? Because you have to have things like plural and singular and verb and noun and adjective. And there’s a lot of things we have to get to before we get to Italy and Japan, and there’s a lot of countries in the world.

Chris Olah4:41:18

那么,模型如何才能在‘线性表征假说’(linear representation hypothesis)成立的前提下,同时表征出多于其维度数量的概念呢?这意味着什么?好吧,既然线性表征假说成立,那必然存在某种有趣的现象正在发生。在进入下一话题前,我再补充一个有趣的现象:之前我们提到过‘多义神经元’(polysemantic neurons)——比如在Inception V1模型中观察到的那些表现优异的神经元,像‘汽车检测器’‘曲线检测器’等,它们会对大量高度一致的刺激产生响应;但同时也存在大量神经元,其响应对象彼此毫不相关,这本身也是一个有趣的现象。此外我们还发现,即便是那些表现极其‘干净’的神经元,若考察其微弱激活状态(例如激活强度仅为最大激活值的5%),此时的响应就未必反映其核心功能了。

英文原文

And so how might it be that models could simultaneously have the linear representation hypothesis be true and also represent more things than they have directions? So what does that mean? Well, okay, so if linear representation hypothesis is true, something interesting has to be going on. Now, I’ll tell you one more interesting thing before we go, and we do that, which is earlier we were talking about all these polysemantic neurons, these neurons that when we were looking at inception V1, these nice neurons that the car detector and the curve detector and so on that respond to lots of very coherent things. But it’s lots of neurons that respond to a bunch of unrelated things. And that’s also an interesting phenomenon. And it turns out as well that even these neurons that are really, really clean, if you look at the weak activations, so if you look at the activations where it’s activating 5% of the maximum activation, it’s really not the core thing that it’s expecting.

Chris Olah4:42:14

举例来说,若观察一个‘曲线检测器’,并聚焦于其激活强度仅为5%的那些情形,你既可将其解读为噪声,也可能意味着它在此时正执行着其他功能。明白吗?那么这究竟是如何实现的呢?数学中存在一个惊人现象,称为‘压缩感知’(compressed sensing)——这是一个令人惊讶的事实:当你将高维空间中的向量投影至低维空间时,通常无法逆向还原出原始高维向量,因为你已丢失了信息。这就像无法对矩形矩阵求逆,只有方阵才可逆。但事实证明,这并非绝对成立:若我告诉你该高维向量是稀疏的(即大部分分量为零),那么你往往能以极高概率成功重构出原始高维向量。

英文原文

So if you look at a curve detector for instance, and you look at the places where it’s 5% active, you could interpret it just as noise or it could be that it’s doing something else there. Okay? So how could that be? Well, there’s this amazing thing in mathematics called compressed sensing, and it’s actually this very surprising fact where if you have a high dimensional space and you project it into a low dimensional space, ordinarily you can’t go and sort of un-projected and get back your high dimensional vector, you threw information away. This is like you can’t invert a rectangular matrix. You can only invert square matrices. But it turns out that that’s actually not quite true. If I tell you that the high-dimensional vector was sparse, so it’s mostly zeros, then it turns out that you can often go and find back the high-dimensional vector with very high probability.

Chris Olah4:43:12

这确实是个令人惊讶的事实,对吧?它表明:只要向量足够稀疏,你就可以将高维向量空间投影到低维空间,且这种降维操作依然有效。‘超叠加假说’(superposition hypothesis)正是主张:神经网络(例如词嵌入)内部发生的正是此类现象。词嵌入既能将特定方向作为有意义的概念载体,又能借助其运行于高维空间这一特性——加之这些概念本身具有稀疏性(例如你通常不会同时提及‘日本’和‘意大利’)——在绝大多数情况下,‘日本’与‘意大利’的表征值均为零,即二者均未出现。若此假设成立,则实际可承载的有意义方向(即特征)数量,便可远超其维度总数。

英文原文

So that’s a surprising fact, right? It says that you can have this high-dimensional vector space, and as long as things are sparse, you can project it down, you can have a lower-dimensional projection of it, and that works. So the superstition hypothesis is saying that that’s what’s going on in neural networks, for instance, that’s what’s going on in word embeddings. The word embeddings are able to simultaneously have directions be the meaningful thing, and by exploiting the fact that they’re operating on a fairly high-dimensional space, they’re actually… and the fact that these concepts are sparse, you usually aren’t talking about Japan and Italy at the same time. Most of those concepts, in most instances, Japan and Italy are both zero. They’re not present at all. And if that’s true, then you can go and have it be the case that you can have many more of these sort of directions that are meaningful, these features than you have dimensions.

Chris Olah4:44:04

同理,在讨论神经元时,你所能表征的概念数量也可远超神经元总数。以上便是‘超叠加假说’的宏观图景。该假说甚至引申出一个更为激进的推论:不仅神经网络的表征方式如此,其计算过程亦可能遵循相同逻辑——包括所有神经元之间的连接关系。某种意义上,神经网络或许是更大规模、更稀疏的神经网络投射出的‘影子’,而我们所观测到的,正是这些投影。‘超叠加假说’最强版本则严肃对待这一观点,进而提出:在更高层面上,确实存在一个‘楼上模型’(upstairs model),其中神经元本身极度稀疏且彼此交织,神经元间的权重构成极其稀疏的电路结构——而这才是我们真正研究的对象;我们所观察到的,不过是该‘楼上模型’投射出的证据之影。我们必须找到那个原始对象。

英文原文

And similarly, when we’re talking about neurons, you can have many more concepts than you have neurons. So that’s at a high level, the superstition hypothesis. Now it has this even wilder implication, which is to go and say that neural networks, it may not just be the case that the representations are like this, but the computation may also be like this. The connections between all of them. And so in some sense, neural networks may be shadows of much larger sparser neural networks. And what we see are these projections. And the strongest version of superstition hypothesis would be to take that really seriously and sort of say there actually is in some sense this upstairs model where the neurons are really sparse and all interpleural, and the weights between them are these really sparse circuits. And that’s what we’re studying. And the thing that we’re observing is the shadow of evidence. We need to find the original object.

Lex Fridman4:45:03

而学习过程,本质上就是在尝试构建一种对‘楼上模型’的压缩表示,使其在投影过程中不致丢失过多信息。

英文原文

And the process of learning is trying to construct a compression of the upstairs model that doesn’t lose too much information in the projection.

Chris Olah4:45:11

是的,这相当于在寻找一种高效拟合方式之类的过程。梯度下降算法正在执行这一任务;事实上,这暗示梯度下降虽可单纯表征一个稠密神经网络,但它本质上是在隐式地搜索一类极其稀疏的模型空间——这些模型经投影后可落入当前低维空间。目前已有大量研究致力于稀疏神经网络:你可以设计出边连接稀疏、激活值也稀疏的神经网络。

英文原文

Yeah, it’s finding how to fit it efficiently or something like this. The gradient descent is doing this and in fact, so this sort of says that gradient descent, it could just represent a dense neural network, but it sort of says that gradient descent is implicitly searching over the space of extremely sparse models that could be projected into this low-dimensional space. And this large body of work of people going and trying to study sparse neural networks where you go and you have… you could design neural networks where the edges are sparse and the activations are sparse.

Chris Olah4:45:38

而我的感觉是,这类工作总体上——它给人的感觉非常有原则性,逻辑上也极其自洽;但据我整体印象,这类工作实际上并未取得特别好的成效。我认为一个可能的解释是:神经网络本身在某种意义上已经是稀疏的。你试图主动去实现这种稀疏性,但梯度下降算法却在后台更高效地遍历了整个稀疏模型空间,自主搜寻出最高效的稀疏模型;随后又设法将其优雅地‘折叠’压缩,以便能便捷地在你的GPU上运行——毕竟GPU擅长执行密集矩阵乘法。而这种机制,你根本无法超越。

英文原文

And my sense is that work has generally, it feels very principled, it makes so much sense, and yet that work hasn’t really panned out that well, is my impression broadly. And I think that a potential answer for that is that actually the neural network is already sparse in some sense. You were trying to go and do this. Gradient descent was actually behind the scenes going and searching more efficiently than you could through the space of sparse models and going and learning whatever sparse model was most efficient. And then figuring out how to fold it down nicely to go and run conveniently on your GPU, which does as nice dense matrix multiplies. And that you just can’t beat that.

Lex Fridman4:46:16

你认为一个神经网络里最多能塞进多少个概念?

英文原文

How many concepts do you think can be shoved into a neural network?

Chris Olah4:46:20

这取决于这些概念的稀疏程度。因此,参数数量很可能构成一个上限,因为你仍需拥有实际可打印的权重来连接它们——这是第一个上限。事实上,压缩感知(compressed sensing)和约翰逊–林登斯特劳斯引理(Johnson-Lindenstrauss lemma)等众多优美的理论结果本质上都在告诉你:如果你有一个向量空间,并希望其中的向量近乎正交——而这大概正是你在此处真正需要的。于是你会说:‘好吧,我放弃要求我的概念(即特征)严格正交,但我希望它们彼此干扰尽可能小;因此我只要求它们近乎正交。’

英文原文

Depends on how sparse they are. So there’s probably an upper bound from the number of parameters because you still have to have print weights that go and connect them together. So that’s one upper bound. There are in fact all these lovely results from compressed sensing and the Johnson-Lindenstrauss lemma and things like this that they basically tell you that if you have a vector space and you want to have almost orthogonal vectors, which is sort of probably the thing that you want here. So you’re going to say, “Well, I’m going to give up on having my concepts, my features be strictly orthogonal, but I’d like them to not interfere that much. I’m going to have to ask them to be almost orthogonal.”

Chris Olah4:46:56

那么这就意味着:一旦你设定好所能容忍的余弦相似度阈值,理论上可容纳的概念数实际上随神经元数量呈指数级增长。因此,在某个阶段,这甚至将不再成为限制因素;而相关理论结果本身已足够优美。事实上,情况或许比这还要更好一些,因为上述结论大致对应的是‘任意一组随机特征都可能被激活’这一假设;但实际上,这些特征之间存在某种相关性结构——某些特征更倾向于共同出现,而另一些则较少共现。因此,依我推测,神经网络在特征打包方面表现极佳,其容量上限很可能并非此处真正的瓶颈。

英文原文

Then this would say that it’s actually for, once you set a threshold for what you’re willing to accept in terms of how much cosine similarity there is, that’s actually exponential in the number of neurons that you have. So at some point, that’s not going to even be the limiting factor, but there’s some beautiful results there. And in fact, it’s probably even better than that in some sense because that’s sort of for saying that any random set of features could be active. But in fact the features have sort of a correlational structure where some features are more likely to co-occur and other ones are less likely to co-occur. And so neural networks, my guest would be, could do very well in terms of going and packing things to the point that’s probably not the limiting factor.

Lex Fridman4:47:37

那么,多义性(polysemanticity)问题在此处如何介入?

英文原文

How does the problem of polysemanticity enter the picture here?

Chris Olah4:47:41

多义性是我们观察到的一种现象:当你审视大量神经元时,会发现单个神经元并不只表征一个概念,它并非一个干净利落的特征;相反,它会对许多互不相关的刺激产生响应。而‘叠加假说’(superposition hypothesis)可被视作一种用于解释多义性观测现象的假说。换言之,多义性是一种已被观察到的现象,而叠加假说则是用以解释该现象(以及一些其他现象)的一种假说。

英文原文

Polysemanticity is this phenomenon we observe where you look at many neurons and the neuron doesn’t just sort of represent one concept, it’s not a clean feature. It responds to a bunch of unrelated things. And superstition you can think of as being a hypothesis that explains the observation of polysemanticity. So polysemanticity is this observed phenomenon and superstition is a hypothesis that would explain it along with some other things.

Lex Fridman4:48:05

因此,这使得机制可解释性研究(Mechinterp)变得更加困难。

英文原文

So that makes Mechinterp more difficult.

Chris Olah4:48:08

没错。如果你试图基于单个神经元来理解模型行为,而这些神经元本身又是多义的,那你就会陷入极大困境。最直接的回答或许是:‘好吧,你正在观察神经元,试图理解它们;但这个神经元对大量事物都有响应,它没有清晰明确的含义——这显然很糟糕。’另一种你可以追问的是:我们最终真正想理解的是权重。假如你有两个多义神经元,每个神经元各自响应三类事物,且二者之间存在一个权重连接,那这个权重究竟意味着什么?是否意味着全部九种组合关系都在同时发生?

英文原文

Right. So if you’re trying to understand things in terms of individual neurons and you have polysemantic neurons, you’re in an awful lot of trouble. The easiest answer is like, “Okay, well you’re looking at the neurons, you’re trying to understand them. This one responds for a lot of things. It doesn’t have a nice meaning. Okay, that’s bad.” Another thing you could ask is ultimately we want to understand the weights. And if you have two polysemantic neurons and each one responds to three things and then the other neuron responds to three things and you have a wait between them, what does that mean? Does it mean that all three, there’s these nine interactions going on?

Chris Olah4:48:40

这确实是一件非常奇怪的事;但背后还存在一个更深层的原因,即神经网络是在极高维空间中运作的。我之前提到过,我们的目标是理解神经网络及其内部机制。有人或许会说:‘它不过是一个数学函数而已,为什么不直接去看它呢?’我最早开展的一个项目就研究了这类将二维空间映射到二维空间的神经网络,你可以用一种极为优美的方式将其诠释为‘弯曲流形’。那我们为何不能沿用这种方式?原因在于:随着输入维度升高,该空间的‘体积’在某种意义上随输入数量呈指数级增长;因此你根本无法对其进行可视化。

英文原文

It’s a very weird thing, but there’s also a deeper reason, which is related to the fact that neural networks operate on really high dimensional spaces. So I said that our goal was to understand neural networks and understand the mechanisms. And one thing you might say is, “Well, it’s just a mathematical function. Why not just look at it, right?” One of the earliest projects I did studied these neural networks that mapped two-dimensional spaces to two-dimensional spaces, and you can sort of interpret them in this beautiful way is like bending manifolds. Why can’t we do that? Well, as you have a higher dimensional space, the volume of that space in some sense is exponential in the number of inputs you have. And so you can’t just go and visualize it.

Chris Olah4:49:19

所以我们必须以某种方式将其拆解——必须将这个指数级增长的空间分解为若干部分,即分解为数量非指数级的一组对象,从而让我们能够独立地进行推理。而‘独立性’至关重要,因为正是这种独立性使我们无需考虑所有指数级组合的可能性。若特征是单义的(monosomatic)、仅具单一含义、确有明确含义,这恰恰是使我们得以独立思考它们的关键前提。因此,若要探寻我们为何追求可解释的单义特征的最深层原因,我认为这才是真正的根本原因。

英文原文

So we somehow need to break that apart. We need to somehow break that exponential space into a bunch of things, some non-exponential number of things that we can reason about independently. And the independence is crucial because it’s the independence that allows you to not have to think about all the exponential combinations of things. And things being monosomatic, things only having one meaning, things having a meaning, that is the key thing that allows you to think about them independently. And so I think if you want the deepest reason why we want to have interpretable monosomatic features, I think that’s really the deep reason.

Lex Fridman4:49:58

因此,你近期工作的目标正是:如何从一个充斥着多义特征及各种混乱现象的神经网络中,提取出单义特征。

英文原文

And so the goal here as your recent work has been aiming at is how do we extract the monosomatic features from a neural net that has polysemantic features and all this mess.

Chris Olah4:50:10

是的,我们观察到了这些多义神经元,并假设其背后机制正是叠加(superposition)。倘若叠加确实是真实机制,那么其实存在一种已有成熟基础的技术,即字典学习(dictionary learning),它正是原理上最恰当的做法。事实证明,若采用字典学习——尤其是以某种方式高效实施、并在某种程度上自然实现良好正则化的稀疏自编码器(sparse autoencoder)——那么,那些原本并不存在的、高度可解释的特征便会自然而然地浮现出来。这并非人们事先必然能预测到的结果,但它的确效果极佳。在我看来,这似乎是对线性表征与叠加假说的一种非平凡验证。

英文原文

Yes, we observe these polysemantic neurons, we hypothesize that’s what’s going on is superposition. And if superposition is what’s going on, there’s actually a sort of well-established technique that is sort of the principled thing to do, which is dictionary learning. And it turns out if you do dictionary learning in particular, if you do sort of a nice efficient way that in some sense sort of nicely regularizes that as well called a sparse auto encoder. If you train a sparse auto encoder, these beautiful interpretable features start to just fall out where there weren’t any beforehand. So that’s not a thing that you would necessarily predict, but it turns out that works very, very well. To me, that seems like some non-trivial validation of linear representations and superposition.

Lex Fridman4:50:51

因此,通过字典学习,你并非在寻找特定类型的类别;你并不预先知道它们是什么,它们只是自行涌现出来。

英文原文

So with dictionary learning, you’re not looking for particular kind of categories. You don’t know what they are, they just emerge.

Chris Olah4:50:57

完全正确。这也呼应了我们早先提出的观点:我们不做任何预设。梯度下降比我们更聪明,所以我们不预设任何存在形式。当然,你完全可以这么做——比如预设存在一个PHP特征,然后专门去搜索它;但我们并未如此行事。我们声明:我们并不知道那里究竟会有什么;相反,我们只是让稀疏自编码器自行去发现那里实际存在的东西。

英文原文

Exactly. And this gets back to our earlier point when we’re not making assumptions. Gradient descent is smarter than us, so we’re not making assumptions about what’s there. I mean, one certainly could do that, right? One could assume that there’s a PHP feature and go and search for it, but we’re not doing that. We’re saying we don’t know what’s going to be there. Instead, we’re just going to go and let the sparse auto encoder discover the things that are there.

Lex Fridman4:51:16

那么,能否谈谈你们去年十月发表的关于单义性(monosematicity)的论文?我听说其中包含许多令人振奋的突破性成果。

英文原文

So can you talk toward monosematicity paper from October last year? I heard a lot of nice breakthrough results.

Chris Olah4:51:24

您这样评价真是太过奖了。是的,这篇论文确实是我们首次成功运用稀疏自编码器的实践。我们选取了一个单层模型,结果发现:若对其实施字典学习,便能发掘出大量真正优质且高度可解释的特征。例如阿拉伯语特征、希伯来语特征、Base64特征等,我们都进行了深入细致的研究,并充分证实了它们确实如我们所预期的那样发挥作用。进一步发现,若训练两个规模加倍的模型(或分别训练两个不同模型),再对它们各自进行字典学习,也能在两个模型中均找到彼此对应的类似特征——这非常有趣。此外,我们还发现了种类繁多的各类特征。因此,这项工作本质上只是有力证明了该方法切实可行。我需要补充说明的是,Cunningham等人几乎在同一时期也取得了非常相似的结果。

英文原文

That’s very kind of you to describe it that way. Yeah, I mean, this was our first real success using sparse autoencoders. So we took a one-layer model, and it turns out if you go and you do dictionary learning on it, you find all these really nice interpretable features. So the Arabic feature, the Hebrew feature, the Base64 features were some examples that we studied in a lot of depth and really showed that they were what we thought they were. Turns out if you train a model twice as well and train two different models and do dictionary learning, you find analogous features in both of them. So that’s fun. You find all kinds of different features. So that was really just showing that this works. And I should mention that there was this Cunningham and all that had very similar results around the same time.

Lex Fridman4:52:08

开展这类小规模实验并发现其实际奏效,本身就充满乐趣。

英文原文

There’s something fun about doing these kinds of small scale experiments and finding that it’s actually working.

Chris Olah4:52:14

是啊,而且这里竟蕴含着如此丰富的结构。或许我们可以稍作回溯:此前一段时间,我一度以为,所有这些机制可解释性研究的最终结果,或许只会得出一个解释——即这类研究本质上异常艰难、不可行。我们会得出结论:‘嗯,这里存在叠加问题,而叠加问题确实极难解决,我们恐怕束手无策了。’但事实并非如此。恰恰相反,一种非常自然、简单的技术居然直接奏效了。因此,这实际上是一种极佳的局面。我认为这是一个难度很高的研究课题,具有大量研究风险,未来仍极有可能失败;但当该方法开始奏效之时,我们已然成功规避了相当一部分重大的研究风险。

英文原文

Yeah, well, and that there’s so much structure here. So maybe stepping back, for a while I thought that maybe all this mechanistic interpolate work, the end result was going to be that I would have an explanation for why it was sort of very hard and not going to be tractable. We’d be like, “Well, there’s this problem with supersession and it turns out supersession is really hard and we’re kind of screwed, but that’s not what happened. In fact, a very natural simple technique just works. And so then that’s actually a very good situation. I think this is a sort of hard research problem and it’s got a lot of research risk and it might still very well fail, but I think that some very significant amount of research risk was put behind us when that started to work.

Lex Fridman4:52:57

那么,能否具体描述一下,通过这种方式可以提取出哪些类型的特征?

英文原文

Can you describe what kind of features can be extracted in this way?

Chris Olah4:53:02

嗯,这取决于你所研究的模型。模型越大,其能力就越复杂、越精细。我们稍后可能还会讨论后续工作。但在这些单层模型中,一些非常常见的现象,我认为包括语言——既涵盖编程语言,也涵盖自然语言。其中存在大量与特定语境下特定词汇相关的特征。例如,‘the’后面很可能紧跟着一个名词,因此你可以将此视为一个特征;但你也可以将其理解为保护某个特定名词特征的机制。此外,还存在一类特征,会在特定语境(比如法律文件或数学文档等)中被激活。例如,在数学语境中,‘the’之后可能预测‘vector’或‘matrix’等数学术语;而在其他语境中,则会预测其他内容——这类现象十分普遍。

英文原文

Well, so it depends on the model that you’re studying. So the larger the model, the more sophisticated they’re going to be. And we’ll probably talk about follow up work in a minute. But in these one layer models, so some very common things I think were languages, both programming languages and natural languages. There were a lot of features that were specific words in specific contexts, so the. And I think really the way to think about this is that the is likely about to be followed by a noun. So you could think of this as the feature, but you could also think of this as protecting a specific noun feature. And there would be these features that would fire for the in the context of say, a legal document or a mathematical document or something like this. And so maybe in the context of math, you’re like the, and then predict vector or matrix, all these mathematical words, whereas in other contexts you would predict other things, that was common.

Lex Fridman4:53:54

本质上,我们需要聪明的人类来为我们所观察到的现象打上标签。

英文原文

And basically we need clever humans to assign labels to what we’re seeing.

Chris Olah4:54:00

是的。这项工作唯一的作用,就是为你展开那些原本折叠在一起的内容。如果所有内容都像被层层叠压一样折叠在自身之上——序列化过程正是将一切堆叠覆盖于其上,导致你根本无法看清细节——那么这项工作就是在将其展开。但即便如此,你面对的仍是一个极其复杂的对象,需要投入大量精力去理解这些展开后的成分;其中有些特征甚至极为微妙。即便是这个单层模型中,关于Unicode也存在一些非常有趣的现象:当然,某些语言采用Unicode编码,而分词器未必为每个Unicode字符都分配一个专属词元(token)。因此,实际出现的往往是交替出现的词元模式,每个词元分别代表某个Unicode字符的一半。

英文原文

Yes. So the only thing this is doing is that sort of unfolding things for you. So if everything was sort of folded over top of it, serialization folded everything on top of itself and you can’t really see it, this is unfolding it. But now you still have a very complex thing to try to understand. So then you have to do a bunch of work understanding what these are, and some are really subtle. There’s some really cool things even in this one layer model about Unicode, where of course some languages are in Unicode, and the tokenizer won’t necessarily have a dedicated token for every Unicode character. So instead, what you’ll have is you’ll have these patterns of alternating token or alternating tokens that each represent half of a Unicode character.

Chris Olah4:54:40

接着,你会有一个不同的特征专门负责在对应位置被激活,起到类似‘好,我刚完成一个字符,请预测下一个前缀’,以及‘好,我现在处于前缀位置,请预测一个合理的后缀’的作用——你需要在这两者之间来回切换。因此,这类交换层(swap layer)模型确实非常有趣。另外,你或许会想:‘应该只有一个Base64特征吧?’但事实却是,实际上存在大量Base64特征,因为英文文本经Base64编码后所产生的Base64词元分布,与常规Base64词元分布截然不同。此外,分词方式本身也存在可被利用的特性。总之,这里面有各种各样的有趣现象。

英文原文

And you have a different feature that goes and activates on the opposing ones to be like, “Okay, I just finished a character, go and predict next prefix. Then okay, I’m on the prefix, predict a reasonable suffix.” And you have to alternate back and forth. So these swap layer models are really interesting. And I mean there’s another thing that you might think, “Okay, there would just be one Base64 feature, but it turns out there’s actually a bunch of Base64 features because you can have English text encoded as Base64, and that has a very different distribution of Base64 tokens than regular. And there’s some things about tokenization as well that it can exploit. And I don’t know, there’s all kinds of fun stuff.

Lex Fridman4:55:21

为当前现象打标签这项任务难度如何?能否由AI自动完成?

英文原文

How difficult is the task of assigning labels to what’s going on? Can this be automated by AI?

Chris Olah4:55:28

嗯,我认为这取决于具体特征,也取决于你对AI的信任程度。目前已有大量工作致力于自动化可解释性(automated interpretability),我认为这是个极具前景的方向;我们自身也开展了相当多的自动化可解释性工作,并让Claude去为我们的特征打标签。

英文原文

Well, I think it depends on the feature, and it also depends on how much you trust your AI. So there’s a lot of work doing automated interoperability. I think that’s a really exciting direction, and we do a fair amount of automated interoperability and have Claude go and label our features.

Lex Fridman4:55:42

有没有一些特别有趣的情形——比如它完全正确,或者完全错误?

英文原文

Is there some fun moments where it’s totally right or it’s totally wrong?

Chris Olah4:55:47

有啊。我觉得很常见的情况是,它给出的描述非常笼统,从某种意义上说虽属实,却并未真正捕捉到正在发生的具体现象。我认为这种情况相当普遍。至于特别有趣的例子,我一时倒想不出一个特别逗的。

英文原文

Yeah, well, I think it’s very common that it says something very general, which is true in some sense, but not really picking up on the specific of what’s going on. So I think that’s a pretty common situation. You don’t know that I have a particularly amusing one.

Lex Fridman4:56:06

这很有意思——那种‘说法没错,却未能触及事物深层细微之处’的微妙落差。这本身就是一个普遍性挑战:AI已能惊人地准确说出正确的话,但有时却缺乏深度。在此背景下,这就像是ARC挑战(一种类似智商测试的任务),而弄清某个特征究竟表征什么,就像解一道需要动脑筋的小谜题。

英文原文

That’s interesting. That little gap between it is true, but it doesn’t quite get to the deep nuance of a thing. That’s a general challenge, it’s already an incredible caution that can say a true thing, but it’s missing the depth sometimes. And in this context, it’s like the ARC challenge, the sort of IQ type of tests. It feels like figuring out what a feature represents is a little puzzle you have to solve.

Chris Olah4:56:35

是的。而且我认为,有些特征相对容易理解,有些则更难。没错,这确实棘手。还有另一点——我不确定这是否部分源于我个人的审美倾向,但我尽量给出理性化的解释:我其实对自动化可解释性略持怀疑态度,部分原因在于,我真心希望人类能够真正理解神经网络。倘若神经网络替我完成了理解,我内心多少有些抵触;不过我也承认……某种程度上,我有点像那些数学家:‘如果证明是由计算机自动生成的,那就不算数’——毕竟人无法真正理解它。此外,我还想到一种关于‘信任之信任’(trusting trust)的反思:有一场著名演讲曾指出,当你编写计算机程序时,你必须信任你的编译器。

英文原文

Yeah. And I think that sometimes they’re easier and sometimes they’re harder as well. Yeah, I think that’s tricky. There’s another thing which I don’t know, maybe in some ways this is my aesthetic coming in, but I’ll try to give you a rationalization. I’m actually a little suspicious of automated interoperability, and I think that partly just that I want humans to understand neural networks. And if the neural network is understanding it for me, I don’t quite like that, but I do have a bit of… In some ways, I’m sort of like the mathematicians who are like, “If there’s a computer automated proof, it doesn’t count.” They won’t understand it. But I do also think that there is this kind of reflections on trusting trust type issue where there’s this famous talk about when you’re writing a computer program, you have to trust your compiler.

Chris Olah4:57:20

倘若你的编译器中藏有恶意软件,它就可能向下一个编译器注入恶意代码,那你可就麻烦了,对吧?同理,若你用神经网络去验证自身神经网络的安全性,那么你所依赖的核心假设便是:‘好吧,该神经网络或许并不安全,你得担心它是否正以某种方式暗中干扰你?’目前这尚不构成重大担忧,但我确实在想:长远来看,如果我们不得不依赖真正强大的AI系统来审计自身的AI系统,这种做法真的值得信赖吗?不过,也许我只是在为自己的立场找借口——归根结底,我只是希望人类最终能彻底理解一切。

英文原文

And if there was malware in your compiler, then it could go and inject malware into the next compiler and you’d be kind of in trouble, right? Well, if you’re using neural networks to go and verify that your neural networks are safe, the hypothesis that you’re trusting for is like, “Okay, well the neural network maybe isn’t safe and you have to worry about is there some way that it could be screwing with you? I think that’s not a big concern now, but I do wonder in the long run, if we have to use really powerful AI systems to go and audit our AI systems, is that actually something we can trust? But maybe I’m just rationalizing because I just want us to have to get to a point where humans understand everything.

Lex Fridman4:57:58

是啊,这简直妙极了,尤其当我们正讨论AI安全性,并试图寻找与AI安全性相关的关键特征(如欺骗行为等)时。那么,让我们谈谈2024年5月发布的《缩放单义性》(Scaling Monosematicity)论文。好的,那么将该方法扩展至Claude 3 Sonnet模型,究竟需要哪些条件?

英文原文

Yeah, I mean that’s hilarious, especially as we talk about AI safety and looking for features that would be relevant to AI safety, like deception and so on. So let’s talk about the Scaling Monosematicity paper in May 2024. Okay. So what did it take to scale this, to apply to Claude 3 Sonnet?

Chris Olah4:58:18

嗯,需要大量GPU。

英文原文

Well, a lot of GPUs.

Lex Fridman4:58:19

更多GPU。明白了。

英文原文

A lot more GPUs. Got it.

Chris Olah4:58:21

不过,我的一位同事汤姆·亨尼根(Tom Henighan)曾参与早期的缩放定律(scaling laws)研究,他从一开始就对‘可解释性是否存在缩放定律’这一问题颇感兴趣。因此,当这项工作初见成效、稀疏自编码器(sparse autoencoders)开始奏效时,他立刻着手探究:扩大稀疏自编码器规模的缩放定律是什么?这又如何与扩大基础模型规模相关联?结果发现,这套方法效果极佳;你甚至可以据此推算:若训练一个给定规模的稀疏自编码器,应使用多少token进行训练等等。这实际上对我们大幅扩展此项工作提供了极大帮助,也使我们得以更轻松地训练超大规模稀疏自编码器——虽然并非训练巨型基础模型本身,但已逐步逼近训练真正超大模型所需的巨大开销。

英文原文

But one of my teammates, Tom Henighan was involved in the original scaling laws work, and something that he was sort of interested in from very early on is are there scaling laws for interoperability? And so something he immediately did when this work started to succeed, and we started to have sparse autoencoders work, was he became very interested in what are the scaling laws for making sparse autoencoders larger and how does that relate to making the base model larger? And so it turns out this works really well and you can use it to sort of project, if you train a sparse autoencoder of a given size, how many tokens should you train on and so on. This was actually a very big help to us in scaling up this work, and made it a lot easier for us to go and train really large sparse autoencoders where it’s not training the big models, but it’s starting to get to a point where it’s actually expensive to go and train the really big ones.

Lex Fridman4:59:21

我的意思是,你得把所有这些任务拆分到大量CPU上运行——

英文原文

I mean, you have to do all this stuff of splitting it across large CPUs-

Chris Olah4:59:26

哦,对,没错!这里同样存在巨大的工程挑战,对吧?是的。因此,一方面存在一个科学问题:如何高效地实现规模化?另一方面,要真正实现规模化,还需投入海量工程工作。你必须精心规划架构,必须对诸多环节进行审慎思考。我很幸运能与一群卓越的工程师共事,因为我本人绝对算不上一名出色的工程师。

英文原文

Oh, yeah. No, I mean there’s a huge engineering challenge here too, right? Yeah. So there’s a scientific question of how do you scale things effectively? And then there’s an enormous amount of engineering to go and scale this up. You have to chart it, you have to think very carefully about a lot of things. I’m lucky to work with a bunch of great engineers because I am definitely not a great engineer.

Lex Fridman4:59:43

尤其是基础设施方面。是的,毫无疑问。总而言之,TLDR(太长不看)版结论就是:它成功了。

英文原文

And the infrastructure especially. Yeah, for sure. So it turns out TLDR, it worked.

Chris Olah4:59:49

成功了。是的。我认为这一点至关重要,因为你本可以设想这样一种世界:你朝着单义性(monospecificity)方向努力,克里斯,这很棒——它在单层模型上奏效了;但单层模型其实非常特殊。也许线性表征假说(linear representation hypothesis)和叠加假说(superposition hypothesis)确实是理解单层模型的正确路径,却未必适用于更大规模的模型。因此,我认为,首先,坎宁安等人(Cunningham et al.)的论文已在一定程度上澄清了这一点,并暗示事实并非如此。

英文原文

It worked. Yeah. And I think this is important because you could have imagined a world where you set after towards monospecificity. Chris, this is great. It works on a one-layer model, but one-layer models are really idiosyncratic. Maybe that’s just something, maybe the linear representation hypothesis and superposition hypothesis is the right way to understand a one-layer model, but it’s not the right way to understand larger models. So I think, I mean, first of all, the Cunningham and all paper sort of cut through that a little bit and sort of suggested that this wasn’t the case.

Chris Olah5:00:18

但《缩放单义性》这篇论文则提供了重要证据,表明即使对于非常大的模型(我们在Claude 3 Sonnet上进行了验证,而当时它已是我们的生产级模型之一),这些模型至少仍能被线性特征显著解释。在它们之上开展字典学习(dictionary learning)是可行的;且随着学习到的特征数量增加,所能解释的现象也越来越多。因此,我认为这是一个相当积极的信号。如今,你还能发现许多真正引人入胜的抽象特征,而且这些特征还是多模态的——即同一概念既能由文本触发,也能由图像触发,这非常有趣。

英文原文

But Scaling Monospecificity sort of I think was significant evidence that even for very large models, and we did it on Claude 3 Sonnet, which at that point was one of our production models. Even these models seemed to be substantially explained, at least by linear features. And doing dictionary learning on them works, and as you learn more features, you go and you explain more and more. So that’s, I think, quite a promising sign. And you find now really fascinating abstract features, and the features are also multimodal. They respond to images and texts for the same concept, which is fun.

Lex Fridman5:00:54

是的。你能详细解释一下吗?我的意思是,后门之类的情况,其实存在大量可举的例子——

英文原文

Yeah. Can you explain that? I mean, backdoor, there’s just a lot of examples that you can-

Chris Olah5:01:01

是的。那我们或许就从这一点开始吧。先举一个例子:我们发现了一些与安全漏洞及后门代码相关的特征。结果发现,这两者其实是两个不同的特征。也就是说,存在一个‘安全漏洞’特征;如果你强制激活它,Claude 就会开始在代码中编写诸如缓冲区溢出之类的安全漏洞。此外,该特征还会被大量其他内容触发,其中一些在数据集里排名靠前的样例包括‘--disable-ssl’之类明显极不安全的指令。

英文原文

Yeah. So maybe let’s start with that. One example to start, which is we found some features around security vulnerabilities and backdoorsing code. So turns out those are actually two different features. So there’s a security vulnerability feature, and if you force it active, Claude it will start to go and write security vulnerabilities like buffer overflows into code. And also fires for all kinds of things, some of the top data set examples where things like dash dash, disable SSL or something like this, which are sort of obviously really insecure.

Lex Fridman5:01:34

所以目前阶段,或许只是因为所呈现的样例本身如此,才显得这些例子略为表面、略为明显。我猜其设想是:未来它或许能检测出更细微的东西,比如欺骗行为、程序缺陷等这类内容。

英文原文

So at this point, maybe it’s just because the examples are presented that way, it’s kind of surface a little bit more obvious examples. I guess the idea is that down the line it might be able to detect more nuance like deception or bugs or that kind of stuff.

Chris Olah5:01:50

是的。不过,我想先区分两件事:其一是特征或概念本身的复杂性,其二是我们所观察样例的微妙程度。当我们展示数据集中排名最靠前的样例时,那些正是最极端、最易触发该特征激活的样例;但这并不意味着它对更细微的情形就完全不响应。例如,这个‘不安全代码’特征,在响应‘彻底禁用安全机制’这类极其明显的操作时最为强烈,但它同样也会响应缓冲区溢出以及代码中更隐蔽的安全漏洞。所有这些特征都是多模态的——你可以提问:‘哪些图像会激活这一特征?’结果发现,‘安全漏洞’特征会被如下图像激活:人们点击 Chrome 浏览器以绕过某网站警告的画面,例如 SSL 证书可能出错之类的提示画面。

英文原文

Yeah. Well, maybe I want to distinguish two things. So one is the complexity of the feature or the concept, right? And the other is the nuance of how subtle the examples we’re looking at, right?. So when we show the top data set examples, those are the most extreme examples that cause that feature to activate. And so it doesn’t mean that it doesn’t fire for more subtle things. So that insecure code feature, the stuff that it fires most strongly for are these really obvious disable the security type things, but it also fires for buffer overflows and more subtle security vulnerabilities in code. These features are all multimodal. You could ask it like, “What images activate this feature?” And it turns out that the security vulnerability feature activates for images of people clicking on Chrome to go past this website, the SSL certificate might be wrong or something like this.

Chris Olah5:02:55

另一件特别有趣的事是‘代码后门’特征:一旦你激活它,Claude 就会编写一个后门,将你的数据转储到某个端口之类。但你也可以问:‘那么,哪些图像会激活‘后门’特征?’答案是:内置隐藏摄像头的设备。显然,已形成一类专门销售看似无害实则暗藏隐藏摄像头设备的人群,而他们的广告图中甚至就包含这种隐藏摄像头!我猜,这算是后门概念在物理世界中的对应版本。它某种程度上揭示了这些概念的抽象程度之高——我一方面为居然真存在整条产业链在售卖此类设备而感到悲哀,另一方面又颇为欣喜:模型竟将这类图像列为该特征最靠前的视觉样例。

英文原文

Another thing that’s very entertaining is there’s backdoors in code feature, like you activate it, it goes and Claude writes a backdoor that will go and dump your data to port or something. But you can ask, “Okay, what images activate the backdoor feature?” It was devices with hidden cameras in them. So there’s a whole apparently genre of people going and selling devices that look innocuous that have hidden cameras, and they have ads that has this hidden camera in it? And I guess that is the physical version of a backdoor. And so it sort of shows you how abstract these concepts are, and I just thought that was… I’m sort of sad that there’s a whole market of people selling devices like that, but I was kind of delighted that that was the thing that it came up with as the top image examples for the feature.

Lex Fridman5:03:36

是的,这很棒。它是多模态的,几乎可称‘多上下文’的;它对单一概念给出了宽泛而坚实的定义,这很棒。

英文原文

Yeah, it’s nice. It’s multimodal. It’s multi almost context. It’s broad, strong definition of a singular concept. It’s nice.

Chris Olah5:03:44

是的。

英文原文

Yeah.

Lex Fridman5:03:45

对我而言,一个尤其令人感兴趣的特征——特别是就人工智能安全而言——便是‘欺骗’与‘说谎’。这类方法未来或许真能检测模型内部是否存在说谎行为,尤其是当模型变得越来越聪明、越来越强大时。可以想见,超级智能模型若具备欺骗能力,便可能向操作者隐瞒自身真实意图等一切相关信息,这无疑构成重大威胁。那么,你们在检测模型内部‘说谎’行为方面,有哪些发现?

英文原文

To me, one of the really interesting features, especially for AI safety, is deception and lying. And the possibility that these kinds of methods could detect lying in a model, especially get smarter and smarter and smarter. Presumably that’s a big threat over super intelligent model that it can deceive the people operating it as to its intentions or any of that kind of stuff. So what have you learned from detecting lying inside models?

Chris Olah5:04:13

是的,我认为这方面我们尚处于早期阶段,但我们已发现不少与欺骗和说谎相关的特征。其中有一个特征,会在人类撒谎或实施欺骗行为时被触发;而当你强制激活它时,Claude 就会开始对你撒谎。因此我们拥有一个‘欺骗’特征。此外,还有各种其他特征,例如‘隐瞒信息’‘拒绝回答问题’,以及‘攫取权力’‘政变’等等。因此,存在大量与‘诡异行为’相关的特征;一旦强制激活它们,Claude 的行为就会……呈现出我们并不希望看到的类型。

英文原文

Yeah, so I think we’re in some ways in early days for that, we find quite a few features related to deception and lying. There’s one feature where it fires for people lying and being deceptive, and you force it active and Claude starts lying to you. So we have a deception feature. I mean, there’s all kinds of other features about withholding information and not answering questions, features about power seeking and coups and stuff like that. So there’s a lot of features that are kind of related to spooky things, and if you force them active Claude will behave in ways that are… they’re not the kinds of behaviors you want.

Lex Fridman5:04:50

在机制可解释性(Mechinterp)领域,您认为接下来最令人兴奋的发展方向有哪些?

英文原文

What are possible next exciting directions to you in the space of Mechinterp?

Chris Olah5:04:56

嗯,可探索的方向很多。首先,我非常希望能达到这样一个阶段:我们能建立‘捷径’,不仅理解特征本身,更能借此理解模型的整个计算过程。对我而言,这才是这项工作的终极目标。目前已有一些相关工作,我们也发布过几项成果。例如,Sam Marks 发表过一篇论文,就做了类似尝试;此外,也已有一些边缘性研究。但我认为,这方面仍有大量工作待完成。而其中一项令人振奋的方向,与我们称之为‘干扰权重’(interference weights)的挑战密切相关。由于‘迷信式关联’(superstition),若仅简单粗暴地查看哪些特征彼此连接,就可能误判某些权重的存在——这些权重其实并不存在于模型的上层结构中,而仅仅是‘迷信式关联’产生的伪影。这是个技术性挑战。与此相关,另一个令人兴奋的方向是:我们可以把稀疏自编码器(sparse autoencoders)想象成一种‘望远镜’。它让我们得以向外眺望,观测到大量此前不可见的特征;随着我们构建出越来越优秀的稀疏自编码器、不断提升字典学习(dictionary learning)能力,我们就能观测到越来越多的‘恒星’,并能聚焦于越来越微小的‘恒星’。现有大量证据表明,我们迄今所见的‘恒星’仍只占极小一部分;我们的神经网络宇宙中,尚有大量‘物质’尚未被我们观测到。或许,我们永远无法制造出足够精密的‘仪器’来观测它们;又或许,其中某些部分在计算上根本不可行、不可观测。因此,这有点像天文学意义上的‘暗物质’——不过并非现代天文学所指的暗物质,而是更接近早期天文学家面对无法解释的物质时那种困惑状态下的‘暗物质’。因此,我经常思考这种‘暗物质’:我们是否终将观测到它?倘若无法观测,这对人工智能安全意味着什么?倘若神经网络中有相当大一部分对我们始终不可达,又该如何是好?

英文原文

Well, there’s a lot of things. So for one thing, I would really like to get to a point where we have shortcuts where we can really understand not just the features, but then use that to understand the computation of models. That relief for me is the ultimate goal of this. And there’s been some work, we put out a few things. There’s a paper from Sam Marks that does some stuff like this, and there’s been, I’d say some work around the edges here. But I think there’s a lot more to do, and I think that will be a very exciting thing that’s related to a challenge we call interference weights. Where due to superstition, if you just sort of naively look at what features are connected together, there may be some weights that don’t exist in the upstairs model, but are just sort of artifacts of superstition. So that’s a technical challenge Related to that, I think another exciting direction is just you might think of sparse autoencoders as being kind of like a telescope. They allow us to look out and see all these features that are out there, and as we build better and better sparse autoencoders, we better and better at dictionary learning, we see more and more stars. And we zoom in on smaller and smaller stars. There’s a lot of evidence that we’re only still seeing a very small fraction of the stars. There’s a lot of matter in our neural network universe that we can’t observe yet. And it may be that we’ll never be able to have fine enough instruments to observe it, and maybe some of it just isn’t possible, isn’t computationally tractable to observe. So it’s sort of a kind of dark matter in not in maybe the sense of modern astronomy of early astronomy when we didn’t know what this unexplained matter is. And so I think a lot about that dark matter and whether we’ll ever observe it and what that means for safety if we can’t observe it, if some significant fraction of neural networks are not accessible to us.

Chris Olah5:06:56

另一个我反复思考的问题是:归根结底,机制可解释性是一种极为微观的插值方法,它试图以极精细的方式理解事物;但我们真正关心的许多问题却是高度宏观的。我们真正关注的是神经网络的行为表现——这恰恰是我最在意的问题。当然,也存在许多其他更大尺度的问题值得探讨。而采用这种极度微观的方法,其优势在于:我们更容易验证‘这是否成立?’;但劣势在于:它离我们真正关心的问题相距甚远。因此,我们现在面临一座需要攀爬的阶梯。我认为关键问题是:我们能否找到更高层级的抽象概念?能否借助这些抽象概念,从这种极度微观的视角向上跃升,从而理解神经网络?

英文原文

Another question that I think a lot about is at the end of the day, mechanistic interpolation is this very microscopic approach to interpolation. It’s trying to understand things in a very fine-grained way, but a lot of the questions we care about are very macroscopic. We care about these questions about neural network behavior, and I think that’s the thing that I care most about. But there’s lots of other sort of larger-scale questions you might care about. And the nice thing about having a very microscopic approach is it’s maybe easier to ask, is this true? But the downside is its much further from the things we care about. And so we now have this ladder to climb. And I think there’s a question of will we be able to find, are there larger-scale abstractions that we can use to understand neural networks that can we get up from this very microscopic approach?

Lex Fridman5:07:48

是的。您曾将此问题表述为某种‘器官级’问题。

英文原文

Yeah. You’ve written about this as kind of organs question.

Chris Olah5:07:52

没错,正是如此。

英文原文

Yeah, exactly.

Lex Fridman5:07:53

如果我们把可解释性比作神经网络的‘解剖学’,那么当前绝大多数研究线索都集中在观察极其微小的‘血管’,即在微观尺度上研究单个神经元及其连接方式。然而,许多自然提出的问题却无法通过这种微观尺度的方法加以解答。相比之下,生物学解剖学中最突出的抽象概念,往往涉及更大尺度的结构,例如单个器官(如心脏)或整个器官系统(如呼吸系统)。因此,我们不禁要问:人工神经网络中是否存在类似‘呼吸系统’‘心脏’或‘脑区’这样的结构?

英文原文

If we think of interpretability as a kind of anatomy of neural networks, most of the circus threads involve studying tiny little veins looking at the small scale and individual neurons and how they connect. However, there are many natural questions that the small-scale approach doesn’t address. In contrast, the most prominent abstractions and biological anatomy involve larger-scale structures like individual organs, like the heart or entire organ systems like the respiratory system. And so we wonder, is there a respiratory system or heart or brain region of an artificial neural network?

Chris Olah5:08:29

是的,完全正确。而且,若我们回顾科学本身,就会发现:许多科学领域都在多个抽象层级上开展研究。例如在生物学中,既有研究蛋白质与分子的分子生物学,也有细胞生物学、研究组织的组织学、解剖学、动物学,乃至生态学。因此,存在大量不同层级的抽象。物理学亦然:既有研究单个粒子的物理学,又有统计物理学,后者则导出了热力学等理论。因此,我们常常面对多个不同层级的抽象。

英文原文

Yeah, exactly. And I mean, if you think about science, right? A lot of scientific fields investigate things at many level of abstraction. In biology, you have molecular biology studying proteins and molecules and so on, and they have cellular biology, and then you have histology studying tissues, and then you have anatomy, and then you have zoology, and then you have ecology. And so you have many, many levels of abstraction or physics, maybe you have a physics of individual particles, and then statistical physics gives you thermodynamics and things like this. And so you often have different levels of abstraction.

Chris Olah5:09:01

我认为,目前我们所拥有的机制可解释性(如果成功的话),某种程度上类似于神经网络的微生物学;但我们真正想要的,更接近于解剖学。你可能会问:‘为什么我们不能直接抵达那个层面?’ 我认为答案至少在很大程度上是出于迷信——实际上,在未先以恰当方式将微观结构拆解开来、再研究其如何相互连接之前,我们很难看清这种宏观结构。但我仍抱有希望:神经网络内部存在远比特征和电路更为宏大的结构,我们将能构建一个涵盖这些更大尺度实体的故事;随后,你便可以细致深入地研究你所关心的具体部分。

英文原文

And I think that right now we have mechanistic interpretability, if it succeeds, is sort of like a microbiology of neural networks, but we want something more like anatomy. And a question you might ask is, “Why can’t you just go there directly?” And I think the answer is superstition, at least in significant part. It’s that it’s actually very hard to see this macroscopic structure without first sort of breaking down the microscopic structure in the right way and then studying how it connects together. But I’m hopeful that there is going to be something much larger than features and circuits and that we’re going to be able to have a story that involves much bigger things. And then you can sort of study in detail the parts you care about.

Lex Fridman5:09:43

我想,在你们这个领域里,你大概相当于神经网络的心理学家或精神科医生。

英文原文

I suppose, in your biology, like a psychologist or a psychiatrist of a neural network.

Chris Olah5:09:48

我认为最美好的情形是:我们不必让这两个领域各自孤立发展,而是能在二者之间架起一座桥梁,使所有高层级的抽象概念都能牢牢扎根于这一坚实、严谨、理想而言根基牢固的基础之上。

英文原文

And I think that the beautiful thing would be if we could go and rather than having disparate fields for those two things, if you could build a bridge between them, such that you could go and have all of your higher level distractions be grounded very firmly in this very solid, more rigorous, ideally foundation.

Lex Fridman5:10:11

你认为人类大脑(即生物神经网络)与人工神经网络之间有何区别?

英文原文

What do you think is the difference between the human brain, the biological neural network and the artificial neural network?

Chris Olah5:10:17

嗯,神经科学家的工作比我们艰难得多。有时我仅凭自己这份工作远比神经科学家轻松这一点,就感到庆幸不已。我们可以记录所有神经元的活动,而且可任意使用海量数据进行记录;顺便说一句,记录过程中神经元本身不会发生变化。你可以消融特定神经元,可以编辑神经元之间的连接等等,甚至还能撤销这些操作——这简直太棒了。你可以对任一神经元进行干预,强制其激活,并观察后续结果;你清楚知道每个神经元与哪些其他神经元相连。神经科学家们费尽心力想获取连接组(connectome),而我们早已拥有连接组,且规模远超秀丽隐杆线虫(C. elegans)。不仅如此,我们不仅掌握连接组,还明确知道哪些神经元彼此兴奋、哪些彼此抑制——这并非仅知二元连接关系,而是确切知晓连接权重;我们可以计算梯度,也清楚每个神经元在计算上究竟执行何种功能。我不知道……诸如此类的优势不胜枚举。我们相较神经科学家拥有太多优势了。然而即便坐拥全部这些优势,理解神经网络依然极其困难。因此,我有时会想:‘天啊,连我们都觉得如此艰难,那么在神经科学的种种约束条件下,这件事岂非近乎不可能?’ 我不知道。或许部分原因在于,我团队中就有几位神经科学家,也许我心里暗想:‘啊,神经科学家们——或许其中一些人愿意接手一个虽仍极富挑战性、但相对容易些的问题,来投身神经网络研究。待我们在理解神经网络这个‘小池塘’(虽仍极难,但毕竟相对容易)中取得进展后,便可重返真正的生物神经科学领域。’

英文原文

Well, the neuroscientists have a much harder job than us. Sometimes I just count my blessings by how much easier my job is than the neuroscientists. So we can record from all the neurons. We can do that on arbitrary amounts of data. The neurons don’t change while you’re doing that, by the way. You can go and ablate neurons, you can edit the connections and so on, and then you can undo those changes. That’s pretty great. You can intervene on any neuron and force it active and see what happens. You know which neurons are connected to everything. Neuroscientists want to get the connectome, we have the connectome and we have it for much bigger than C. elegans. And then not only do we have the connectome, we know which neurons excite or inhibit each other, right? It’s not just that we know the binary mask, we know the weights. We can take gradients, we know computationally what each neuron does. I don’t know. The list goes on and on. We just have so many advantages over neuroscientists. And then despite having all those advantages, it’s really hard. And so one thing I do sometimes think is like, “Gosh, if it’s this hard for us, it seems impossible under the constraints of neuroscience or near impossible.” I don’t know. Maybe part of me is I’ve got a few neuroscientists on my team, maybe I’m sort of like, “Ah, the neuroscientists. Maybe some of them would like to have an easier problem that’s still very hard, and they could come and work on neural networks. And then after we figure out things in sort of the easy little pond of trying to understand neural networks, which is still very hard, then we could go back to biological neuroscience.”

Lex Fridman5:11:51

我很喜欢你关于机制可解释性(MechInterp)研究目标的阐述——即安全与美这两大目标。那么,能否谈谈‘美’这一面向?

英文原文

I love what you’ve written about the goal of MechInterp research as two goals, safety and beauty. So can you talk about the beauty side of things?

Chris Olah5:11:59

是的。这里有个有趣的现象:我认为某些人对神经网络多少有些失望,他们会觉得:‘啊,神经网络嘛,不过就是些简单规则罢了;然后你只需做大量工程工作将其放大,它就表现得非常出色。可那些复杂的理念在哪儿呢?这算不上什么优美、漂亮的科学成果。’ 有时听人这么说,我脑海中浮现的画面却是:‘进化论真无聊啊——不就一堆简单规则吗?你让它运行漫长时光,结果就得到了生物学。生物学竟以这种方式呈现,真是糟透了!复杂规则在哪儿呢?’ 但真正的美恰恰在于:正是这种简洁催生了复杂。

英文原文

Yeah. So there’s this funny thing where I think some people are kind of disappointed by neural networks, I think, where they’re like, “Ah, neural networks, it’s just these simple rules. Then you just do a bunch of engineering to scale it up and it works really well. And where’s the complex ideas? This isn’t a very nice, beautiful scientific result.” And I sometimes think when people say that, I picture them being like, “Evolution is so boring. It’s just a bunch of simple rules. And you run evolution for a long time and you get biology. What a sucky way for biology to have turned out. Where’s the complex rules?” But the beauty is that the simplicity generates complexity.

Chris Olah5:12:41

生物学也遵循着这些简单规则,却由此孕育出我们周围所见的一切生命与生态系统;大自然的所有壮丽之美,全都源于进化,源于进化中某种极为简单的东西。同样,我认为神经网络也在其内部构建并创造出巨大的复杂性与美感,以及丰富的内在结构;而人们通常并不去观察、也不试图理解这些结构,只因理解起来实在困难。但我相信,神经网络内部蕴藏着极其丰富、亟待发现的结构,只要我们愿意花时间去观察、去理解,就能领略其中深邃无比的美。

英文原文

Biology has these simple rules and it gives rise to all the life and ecosystems that we see around us. All the beauty of nature, that all just comes from evolution and from something very simple in evolution. And similarly, I think that neural networks build, create enormous complexity and beauty inside and structure inside themselves that people generally don’t look at and don’t try to understand because it’s hard to understand. But I think that there is an incredibly rich structure to be discovered inside neural networks, a lot of very deep beauty if we’re just willing to take the time to go and see it and understand it.

Lex Fridman5:13:20

是的,我热爱机制可解释性(Mechinterp)研究。那种我们正在理解、或至少正窥见其内部魔法运作方式的感觉,真的非常美妙。

英文原文

Yeah, I love Mechinterp. The feeling like we are understanding or getting glimpses of understanding the magic that’s going on inside is really wonderful.

Chris Olah5:13:30

对我而言,这似乎是一个呼之欲出、亟待提出的问题——我是说,其实很多人已在思考这个问题,但我常常惊讶于为何没有更多人关注它:我们为何至今尚不知晓如何设计出能完成这些任务的计算机系统?然而,我们却已拥有这些惊人的系统;我们无法直接编写出能完成这些任务的计算机程序,但这些神经网络却能完成所有这些惊人之事。这感觉显然就是那个亟待解答的核心问题。只要你怀有任何程度的好奇心,就会忍不住发问:‘人类如今竟已造出这些能完成此类任务、而我们自身却不知如何实现的器物,这究竟是怎么回事?’

英文原文

It feels to me like one of the questions that’s just calling out to be asked, and I’m sort of, I mean a lot of people are thinking about this, but I’m often surprised that not more are is how is it that we don’t know how to create computer systems that can do these things? And yet we have these amazing systems that we don’t know how to directly create computer programs that can do these things, but these neural networks can do all these amazing things. And it just feels like that is obviously the question that is calling out to be answered. If you have any degree of curiosity, it’s like, “How is it that humanity now has these artifacts that can do these things that we don’t know how to do?”

Lex Fridman5:14:06

是的,我喜欢‘马戏团朝着目标函数之光伸出手去’这个意象。

英文原文

Yeah. I love the image of the circus reaching towards the light of the objective function.

Chris Olah5:14:11

是的,这是一种有机生长出来的事物,而我们对其究竟长成了什么,却一无所知。

英文原文

Yeah, it’s this organic thing that we’ve grown and we have no idea what we’ve grown.

Lex Fridman5:14:15

那么,感谢你为安全性研究付出的努力,感谢你欣赏自己所发现事物的美,也感谢你今天拨冗交谈,Chris,这真是一场精彩的对话。

英文原文

Well, thank you for working on safety, and thank you for appreciating the beauty of the things you discover. And thank you for talking today, Chris, this was wonderful.

Chris Olah5:14:23

也感谢你抽出时间与我交流。

英文原文

Thank you for taking the time to chat as well.

Lex Fridman5:14:26

感谢收听本期与Chris Olah的对话,以及此前与Dario Amodei和Amanda Askell的对话。如欲支持本播客,请查看简介中的赞助商信息。最后,让我以艾伦·瓦茨(Alan Watts)的一段话作结:‘要理解变化的唯一方式,便是纵身跃入其中,随其流动,并加入这场舞蹈。’ 感谢收听,期待下次再见。

英文原文

Thanks for listening to this conversation with Chris Ola and before that, with Dario Amodei and Amanda Askell. To support this podcast, please check out our sponsors in the description. And now let me leave you with some words from Alan Watts. “The only way to make sense out of change is to plunge into it, move with it, and join the dance.” Thank you for listening and hope to see you next time.

出版方介绍

Dario Amodei is the CEO of Anthropic, the company that created Claude. Amanda Askell is an AI researcher working on Claude’s character and personality. Chris Olah is an AI researcher working on mechanistic interpretability.

编辑摘要和译文均基于该出版方提供的资料生成。

发布者音频 (在新标签页中打开)

来源与研究方法

这些观点关联至发布方提供的带时间戳的访谈文字稿。其中的转述内容已标注,不作为逐字引语呈现。

阅读发布者的文字稿 (在新标签页中打开)报告问题