Transformer 之后是什么?

The Data Exchange ·

Zuzanna Stamirowska 将 2023 年和 2024 年的 RAG、上下文与记忆称为瓶颈,她将其描述为“如果你愿意这么说的话,就是没有天花板的人工智能”。她回忆起自己意识到推理不应仅限于语言层面,并表示潜在思维可以帮助用“极小、极小的模型”求解约束满足问题和数学问题。 阅读 3 条观点,查看支持证据与原始来源。

理解这篇

3 个要点

综合解读

  1. RAG、上下文与记忆在2023–2024年间构成瓶颈

    斯塔米罗夫斯卡表示,该团队在2023年和2024年已察觉RAG、上下文与记忆构成瓶颈;这些方法对于她所描述的“若你愿意,即为‘无天花板的人工智能’”而言,既不高效,也不充分。

    支持这项说法 1

    但非常具体地说,促使我们创办 Pathway 并开始研究后 transformer 架构的原因是,我们期待推理成为一件真正的事。我们在处理机器学习中的时间维度方面——也就是任何形式的在线学习——拥有非常深厚的背景。而我们看到 RAG、上下文和记忆,这里说的是 2023 年和 2024 年,成为了瓶颈。它们不够高效,也不足以实现我们所期望的人工智能——如果你愿意这么说的话,就是没有天花板的人工智能。

    Zuzanna Stamirowska · Publisher transcript paragraph 7

    原始摘录
    But very concretely, what prompted us to start Pathway and to start working on post-transformer architectures was that we were looking forward to reasoning becoming a thing. We have a very strong background in dealing with the dimension of time in machine learning—any sort of online learning. And we saw RAG, context, and memory, and we’re talking about 2023 and 2024, as bottlenecks. They were not efficient or sufficient to unlock AI as we would like—AI with no ceiling, if you wish.
    回到原文语境 →
  2. Stamirowska 称推理不应仅限于语言层面

    Stamirowska 表示:“我们意识到,推理不应仅限于语言层面。”她说,在思维链中将事物用语言表达出来并非最优:这相当于“把它粘在 transformer 之上”,而 transformer 天然不擅长处理时间和从经验中学习。

    支持这项说法 1

    具体来说,我们首先讨论的是持续学习、记忆和推理。随着推进,我们意识到推理不应仅限于语言层面。在思维链中将事物用语言表达出来并不是进行推理的最优方式。使用思维链时,我们实质上是在把它粘在 transformer 之上,而 transformer 天然不擅长处理时间和从经验中学习。

    Zuzanna Stamirowska · Publisher transcript paragraph 8

    原始摘录
    Specifically, we’re talking about continual learning, memory, and reasoning first. And as we moved through it, we realized that reasoning shouldn’t just be linguistic. Verbalizing things in chain of thought is not the optimal way to do reasoning. With chain of thought, we’re essentially gluing it on top of a transformer, and transformers are naturally ill-suited to dealing with time and learning from experience.
    回到原文语境 →

    继续探索

    推理模态 →
  3. 潜在思维可助力求解约束满足问题与数学问题

    斯塔米罗夫斯卡表示,潜在思维可助力以‘极小、极小的模型’求解约束满足问题与数学问题。

    支持这项说法 1

    潜在思维具有巨大优势。正如我们所展示的那样,它能帮助我们原生地解决约束满足问题。我们正使用极小、极小的模型来解决约束满足问题和数学问题。此外还有成本效益:潜在推理正带来惊人的成本降低。

    Zuzanna Stamirowska · Publisher transcript paragraph 37

    原始摘录
    Latent thinking has huge benefits. It can help us solve constraint-satisfaction problems natively, like what we’re showing. We’re solving constraint-satisfaction problems and math with tiny, tiny models. And then there’s cost efficiency. There are just insane cost gains coming from latent reasoning.
    回到原文语境 →

关键时刻3

简短、标注来源的段落,并附有可验证上下文。完整对话保留在其发布者处。

后Transformer架构的动因

RAG、上下文与记忆在2023–2024年间构成瓶颈

但非常具体地说,促使我们创办 Pathway 并开始研究后 transformer 架构的原因是,我们期待推理成为一件真正的事。我们在处理机器学习中的时间维度方面——也就是任何形式的在线学习——拥有非常深厚的背景。而我们看到 RAG、上下文和记忆,这里说的是 2023 年和 2024 年,成为了瓶颈。它们不够高效,也不足以实现我们所期望的人工智能——如果你愿意这么说的话,就是没有天花板的人工智能。

原始摘录
But very concretely, what prompted us to start Pathway and to start working on post-transformer architectures was that we were looking forward to reasoning becoming a thing. We have a very strong background in dealing with the dimension of time in machine learning—any sort of online learning. And we saw RAG, context, and memory, and we’re talking about 2023 and 2024, as bottlenecks. They were not efficient or sufficient to unlock AI as we would like—AI with no ceiling, if you wish.
推理模态

Stamirowska 称推理不应仅限于语言层面

具体来说,我们首先讨论的是持续学习、记忆和推理。随着推进,我们意识到推理不应仅限于语言层面。在思维链中将事物用语言表达出来并不是进行推理的最优方式。使用思维链时,我们实质上是在把它粘在 transformer 之上,而 transformer 天然不擅长处理时间和从经验中学习。

原始摘录
Specifically, we’re talking about continual learning, memory, and reasoning first. And as we moved through it, we realized that reasoning shouldn’t just be linguistic. Verbalizing things in chain of thought is not the optimal way to do reasoning. With chain of thought, we’re essentially gluing it on top of a transformer, and transformers are naturally ill-suited to dealing with time and learning from experience.
潜在推理的优势

潜在思维可助力求解约束满足问题与数学问题

潜在思维具有巨大优势。正如我们所展示的那样,它能帮助我们原生地解决约束满足问题。我们正使用极小、极小的模型来解决约束满足问题和数学问题。此外还有成本效益:潜在推理正带来惊人的成本降低。

原始摘录
Latent thinking has huge benefits. It can help us solve constraint-satisfaction problems natively, like what we’re showing. We’re solving constraint-satisfaction problems and math with tiny, tiny models. And then there’s cost efficiency. There are just insane cost gains coming from latent reasoning.

来源与研究方法

这些观点均关联原始来源。转述已明确标注,不作为逐字原话展示。

打开转录或来源材料 (在新标签页中打开)报告问题

继续了解这些人物的观点