通过带 MLX 后端的 K-Search 实现接近专家级的 Apple Silicon 性能
借助新型 CUDA 到 MLX 翻译层的 K-Search,在 Apple Silicon 上的 Mamba SSM 内核上实现了相较 mlx-lm 高达 20× 的 prefill 加速;在原生 MLX Attention 内核上则达到 0.97× 的加速。作者指出,该方法适用于任何可迁移 CUDA 专业知识的生态系统。
支持这项说法
我们证明,我们的方法在 Apple Silicon 上达到了接近专家级的性能,与原生 MLX Attention 内核相比实现了 0.97 倍的加速,并且在 Mamba SSM 内核上与社区的 mlx-lm 实现相比预填充速度最高提升达 20 倍。
原始摘录
We show that our approach reaches near-expert level performance on Apple Silicon with 0.97x speedup compared to the native MLX Attention kernel, and up to a 20x prefill speedup over the community mlx-lm implementation on the Mamba SSM kernel
上下文
;我们在下文中报告了相关数据,以及有多少增益来自转换层。尽管我们关注的是用于 Apple Silicon 的 MLX 内核,但该方法并不局限于 MLX,而是适用于任何可迁移 CUDA 专业知识的生态系统。
原始上下文
; we report the numbers, and how much of the gain comes from the translation layer, in the sections below. Although we focus on MLX kernels for Apple Silicon, the method is not specific to MLX and applies to any ecosystem where CUDA expertise is transferable.