从 CUDA 到 MLX:K-Search 如何将数十年的内核专业知识引入 Apple Silicon

Berkeley BAIR Blog ·

一份关于 K-Search 的技术报告,该方法使用 MLX 和 Metal 将 CUDA 内核移植到 Apple Silicon。报告聚焦于 LLM 引导翻译中的硬件感知约束,指出上下文质量(而非模型能力)是关键瓶颈,报告了相对于原生 MLX 内核的性能,并提到在两个内核(attention 和 Mamba SSM)上观察到的成功,同时承认对更广泛泛化的不确定性。 阅读 4 条观点,查看支持证据与原始来源。

Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez

理解这篇

4 个要点

综合解读

  1. 缺乏硬件感知约束时,简单的 LLM 移植会失败

    提示 LLM 将 CUDA 内核移植到 MLX/Metal 会生成语法有效但架构不正确的代码,除非由深入的硬件上下文和显式约束加以引导。

    支持这项说法 1

    仅将一个 CUDA 内核交给 LLM 并要求其移植是不够的:若缺乏深入的硬件上下文,它生成的代码虽语法正确,但在架构层面是错误的。

    Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez · 段落 26

    原始摘录
    Simply handing an LLM a CUDA kernel and asking it to port it is not enough: without deep hardware context, it produces code that is syntactically valid but architecturally wrong
    上下文

    然而,更具挑战性的问题并不只是在 MLX 上运行 K-Search。关键洞见在于,专家级 CUDA 内核编码了数十年的优化知识,如果你能弥合概念上的鸿沟,这些知识可以迁移到 Apple GPU。(错误的分块大小、无效的原语、不匹配的内存假设)。

    原始上下文

    However, the more interesting challenge was not simply running K-Search on MLX. The key insight is that expert CUDA kernels encode decades of optimization knowledge that is transferable to Apple GPU if you can bridge the conceptual gap. (wrong tile sizes, invalid primitives, mismatched memory assumptions).

    回到原文语境 →
  2. 瓶颈在于上下文质量,而非 LLM 编码能力

    针对新硬件的 AI 驱动内核优化的主要瓶颈并非 LLM 生成 Metal 代码的能力,而是用于引导搜索的结构化跨平台翻译知识的质量。

    支持这项说法 1

    对我们而言,主要结论在于:瓶颈并非大语言模型(LLM)编写 Metal 代码的能力,而在于我们为其提供的上下文与约束条件的质量。

    Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez · 段落 45

    原始摘录
    For us the main takeaway is that the bottleneck was not the LLM’s ability to write Metal code, but the quality of the context and constraints we gave it.
    上下文

    我们的 CUDA 转换层将现有的 NVIDIA 内核专业知识转化为针对 Apple Silicon 的可操作指导,并由 K-Search 的进化式搜索完成其余工作。

    原始上下文

    Our CUDA translation layer converts existing NVIDIA kernel expertise into actionable guidance for Apple Silicon, and lets K-Search’s evolutionary search do the rest.

    回到原文语境 →
  3. 通过带 MLX 后端的 K-Search 实现接近专家级的 Apple Silicon 性能

    借助新型 CUDA 到 MLX 翻译层的 K-Search,在 Apple Silicon 上的 Mamba SSM 内核上实现了相较 mlx-lm 高达 20× 的 prefill 加速;在原生 MLX Attention 内核上则达到 0.97× 的加速。作者指出,该方法适用于任何可迁移 CUDA 专业知识的生态系统。

    支持这项说法 1

    我们证明,我们的方法在 Apple Silicon 上达到了接近专家级的性能,与原生 MLX Attention 内核相比实现了 0.97 倍的加速,并且在 Mamba SSM 内核上与社区的 mlx-lm 实现相比预填充速度最高提升达 20 倍。

    Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez · 段落 5

    原始摘录
    We show that our approach reaches near-expert level performance on Apple Silicon with 0.97x speedup compared to the native MLX Attention kernel, and up to a 20x prefill speedup over the community mlx-lm implementation on the Mamba SSM kernel
    上下文

    ;我们在下文中报告了相关数据,以及有多少增益来自转换层。尽管我们关注的是用于 Apple Silicon 的 MLX 内核,但该方法并不局限于 MLX,而是适用于任何可迁移 CUDA 专业知识的生态系统。

    原始上下文

    ; we report the numbers, and how much of the gain comes from the translation layer, in the sections below. Although we focus on MLX kernels for Apple Silicon, the method is not specific to MLX and applies to any ecosystem where CUDA expertise is transferable.

    回到原文语境 →
  4. AI 驱动的内核搜索在两个内核上达到接近专家级的性能

    在所研究的两个内核(attention 和 Mamba SSM)上,以结构化跨平台翻译为基础的 AI 驱动演化式内核搜索,在 Apple Silicon 上达到了接近专家级的性能,而无需 GPU 专家从头编写。作者表示,他们尚不清楚这种方法的泛化程度。

    支持这项说法 1

    在我们研究的两个内核上,以结构化跨平台翻译知识为根基、由人工智能驱动的进化式内核搜索,在无需 GPU 专家团队从零起步的情况下,即在 Apple Silicon 上实现了接近专家级的性能。目前尚不清楚该方法的泛化能力究竟如何,但这一结果令人鼓舞。

    Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez · 段落 44

    原始摘录
    On the two kernels we studied, AI-driven evolutionary kernel search grounded in structured cross-platform translation knowledge reached near-expert performance on Apple Silicon without a team of GPU experts starting from scratch. We do not yet know how far this generalizes
    上下文

    但该结果令人鼓舞。

    原始上下文

    , but the result is encouraging.

    回到原文语境 →

关键段落4

带明确归属与语境的原文片段。打开原始文本核查出处。

泛化不确定性

AI 驱动的内核搜索在两个内核上达到接近专家级的性能

在我们研究的两个内核上,以结构化跨平台翻译知识为根基、由人工智能驱动的进化式内核搜索,在无需 GPU 专家团队从零起步的情况下,即在 Apple Silicon 上实现了接近专家级的性能。目前尚不清楚该方法的泛化能力究竟如何,但这一结果令人鼓舞。

原始摘录
On the two kernels we studied, AI-driven evolutionary kernel search grounded in structured cross-platform translation knowledge reached near-expert performance on Apple Silicon without a team of GPU experts starting from scratch. We do not yet know how far this generalizes
上下文

但该结果令人鼓舞。

原始上下文

, but the result is encouraging.

K-Search MLX 性能结果

通过带 MLX 后端的 K-Search 实现接近专家级的 Apple Silicon 性能

我们证明,我们的方法在 Apple Silicon 上达到了接近专家级的性能,与原生 MLX Attention 内核相比实现了 0.97 倍的加速,并且在 Mamba SSM 内核上与社区的 mlx-lm 实现相比预填充速度最高提升达 20 倍。

原始摘录
We show that our approach reaches near-expert level performance on Apple Silicon with 0.97x speedup compared to the native MLX Attention kernel, and up to a 20x prefill speedup over the community mlx-lm implementation on the Mamba SSM kernel
上下文

;我们在下文中报告了相关数据,以及有多少增益来自转换层。尽管我们关注的是用于 Apple Silicon 的 MLX 内核,但该方法并不局限于 MLX,而是适用于任何可迁移 CUDA 专业知识的生态系统。

原始上下文

; we report the numbers, and how much of the gain comes from the translation layer, in the sections below. Although we focus on MLX kernels for Apple Silicon, the method is not specific to MLX and applies to any ecosystem where CUDA expertise is transferable.

上下文质量优于模型能力

瓶颈在于上下文质量,而非 LLM 编码能力

对我们而言,主要结论在于:瓶颈并非大语言模型(LLM)编写 Metal 代码的能力,而在于我们为其提供的上下文与约束条件的质量。

原始摘录
For us the main takeaway is that the bottleneck was not the LLM’s ability to write Metal code, but the quality of the context and constraints we gave it.
上下文

我们的 CUDA 转换层将现有的 NVIDIA 内核专业知识转化为针对 Apple Silicon 的可操作指导,并由 K-Search 的进化式搜索完成其余工作。

原始上下文

Our CUDA translation layer converts existing NVIDIA kernel expertise into actionable guidance for Apple Silicon, and lets K-Search’s evolutionary search do the rest.

LLM 引导的内核移植

缺乏硬件感知约束时,简单的 LLM 移植会失败

仅将一个 CUDA 内核交给 LLM 并要求其移植是不够的:若缺乏深入的硬件上下文,它生成的代码虽语法正确,但在架构层面是错误的。

原始摘录
Simply handing an LLM a CUDA kernel and asking it to port it is not enough: without deep hardware context, it produces code that is syntactically valid but architecturally wrong
上下文

然而,更具挑战性的问题并不只是在 MLX 上运行 K-Search。关键洞见在于,专家级 CUDA 内核编码了数十年的优化知识,如果你能弥合概念上的鸿沟,这些知识可以迁移到 Apple GPU。(错误的分块大小、无效的原语、不匹配的内存假设)。

原始上下文

However, the more interesting challenge was not simply running K-Search on MLX. The key insight is that expert CUDA kernels encode decades of optimization knowledge that is transferable to Apple GPU if you can bridge the conceptual gap. (wrong tile sizes, invalid primitives, mismatched memory assumptions).

这里提到的

全部提及对象

MLX

仅提及

作者报告了其方法在 Apple Silicon 上 MLX 内核(Attention 和 Mamba SSM)的性能结果,将 MLX 视为目标运行环境,而非被批评或背书的工具。

查看支持证据 · Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez

来源与研究方法

这些观点均关联原始来源。转述已明确标注,不作为逐字原话展示。

打开转录或来源材料 (在新标签页中打开)报告问题