作者报告了其方法在 Apple Silicon 上 MLX 内核(Attention 和 Mamba SSM)的性能结果,将 MLX 视为目标运行环境,而非被批评或背书的工具。
Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez ·
支持这项说法
从 CUDA 到 MLX:K-Search 如何将数十年的内核专业知识引入 Apple Silicon
我们证明,我们的方法在 Apple Silicon 上达到了接近专家级的性能,与原生 MLX Attention 内核相比实现了 0.97 倍的加速,并且在 Mamba SSM 内核上与社区的 mlx-lm 实现相比预填充速度最高提升达 20 倍。
原始摘录
We show that our approach reaches near-expert level performance on Apple Silicon with 0.97x speedup compared to the native MLX Attention kernel, and up to a 20x prefill speedup over the community mlx-lm implementation on the Mamba SSM kernel
上下文
;我们在下文中报告了相关数据,以及有多少增益来自转换层。尽管我们关注的是用于 Apple Silicon 的 MLX 内核,但该方法并不局限于 MLX,而是适用于任何可迁移 CUDA 专业知识的生态系统。
原始上下文
; we report the numbers, and how much of the gain comes from the translation layer, in the sections below. Although we focus on MLX kernels for Apple Silicon, the method is not specific to MLX and applies to any ecosystem where CUDA expertise is transferable.