From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon

Berkeley BAIR Blog ·

A technical report on K-Search, a method for porting CUDA kernels to Apple Silicon using MLX and Metal. It focuses on hardware-aware constraints in LLM-guided translation, identifies context quality—not model capability—as the key bottleneck, reports performance relative to native MLX kernels, and notes observed success on two kernels (attention and Mamba SSM), while acknowledging uncertainty about broader generalization. Lies 4 Standpunkte mit Belegen und Links zu den Originalquellen.

Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez

Auf einen Blick

  • Naive LLM porting fails without hardware-aware constraints

    Prompting an LLM to port CUDA kernels to MLX/Metal produces syntactically valid but architecturally incorrect code unless guided by deep hardware context and explicit constraints.

    Unterstützendes Moment lesen · Absatz 26
  • Bottleneck is context quality, not LLM coding ability

    The main bottleneck in AI-driven kernel optimization for new hardware is not the LLM’s ability to generate Metal code, but the quality of the structured, cross-platform translation knowledge used to guide the search.

    Unterstützendes Moment lesen · Absatz 45
  • Near-expert Apple Silicon performance via K-Search with MLX backend

    K-Search with a novel CUDA-to-MLX translation layer achieved 0.97× speedup versus the native MLX Attention kernel and up to 20× prefill speedup over mlx-lm on the Mamba SSM kernel on Apple Silicon. The authors note the method applies to any ecosystem where CUDA expertise is transferable.

    Unterstützendes Moment lesen · Absatz 5
  • AI-driven kernel search reached near-expert performance on two kernels

    On the two kernels studied—attention and Mamba SSM—AI-driven evolutionary kernel search, grounded in structured cross-platform translation, reached near-expert performance on Apple Silicon without requiring GPU experts to start from scratch. The authors state they do not yet know how far this generalizes.

    Unterstützendes Moment lesen · Absatz 44

Wichtige Passagen4

Zugeordnete Passagen mit dem Kontext zur Überprüfung. Öffnen Sie den Originaltext, um die Quelle zu prüfen.

Generalization uncertainty

AI-driven kernel search reached near-expert performance on two kernels

Originalauszug

On the two kernels we studied, AI-driven evolutionary kernel search grounded in structured cross-platform translation knowledge reached near-expert performance on Apple Silicon without a team of GPU experts starting from scratch. We do not yet know how far this generalizes
Kontext

, but the result is encouraging.

K-Search MLX performance results

Near-expert Apple Silicon performance via K-Search with MLX backend

Originalauszug

We show that our approach reaches near-expert level performance on Apple Silicon with 0.97x speedup compared to the native MLX Attention kernel, and up to a 20x prefill speedup over the community mlx-lm implementation on the Mamba SSM kernel
Kontext

; we report the numbers, and how much of the gain comes from the translation layer, in the sections below. Although we focus on MLX kernels for Apple Silicon, the method is not specific to MLX and applies to any ecosystem where CUDA expertise is transferable.

Context quality over model capability

Bottleneck is context quality, not LLM coding ability

Originalauszug

For us the main takeaway is that the bottleneck was not the LLM’s ability to write Metal code, but the quality of the context and constraints we gave it.
Kontext

Our CUDA translation layer converts existing NVIDIA kernel expertise into actionable guidance for Apple Silicon, and lets K-Search’s evolutionary search do the rest.

LLM-guided kernel porting

Naive LLM porting fails without hardware-aware constraints

Originalauszug

Simply handing an LLM a CUDA kernel and asking it to port it is not enough: without deep hardware context, it produces code that is syntactically valid but architecturally wrong
Kontext

However, the more interesting challenge was not simply running K-Search on MLX. The key insight is that expert CUDA kernels encode decades of optimization knowledge that is transferable to Apple GPU if you can bridge the conceptual gap. (wrong tile sizes, invalid primitives, mismatched memory assumptions).

Quelle & Methodik

Diese Standpunkte sind mit ihren Originalquellen verknüpft. Paraphrasen sind gekennzeichnet und keine wörtlichen Zitate.

Transkript oder Quellenmaterial öffnen (wird in einem neuen Tab geöffnet)Ein Problem melden