From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon

Berkeley BAIR Blog ·

A technical report on K-Search, a method for porting CUDA kernels to Apple Silicon using MLX and Metal. It focuses on hardware-aware constraints in LLM-guided translation, identifies context quality—not model capability—as the key bottleneck, reports performance relative to native MLX kernels, and notes observed success on two kernels (attention and Mamba SSM), while acknowledging uncertainty about broader generalization. Lee 4 puntos de vista con sus evidencias y enlaces a las fuentes.

Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez

De un vistazo

  • Naive LLM porting fails without hardware-aware constraints

    Prompting an LLM to port CUDA kernels to MLX/Metal produces syntactically valid but architecturally incorrect code unless guided by deep hardware context and explicit constraints.

    Ver el momento de apoyo · Párrafo 26
  • Bottleneck is context quality, not LLM coding ability

    The main bottleneck in AI-driven kernel optimization for new hardware is not the LLM’s ability to generate Metal code, but the quality of the structured, cross-platform translation knowledge used to guide the search.

    Ver el momento de apoyo · Párrafo 45
  • Near-expert Apple Silicon performance via K-Search with MLX backend

    K-Search with a novel CUDA-to-MLX translation layer achieved 0.97× speedup versus the native MLX Attention kernel and up to 20× prefill speedup over mlx-lm on the Mamba SSM kernel on Apple Silicon. The authors note the method applies to any ecosystem where CUDA expertise is transferable.

    Ver el momento de apoyo · Párrafo 5
  • AI-driven kernel search reached near-expert performance on two kernels

    On the two kernels studied—attention and Mamba SSM—AI-driven evolutionary kernel search, grounded in structured cross-platform translation, reached near-expert performance on Apple Silicon without requiring GPU experts to start from scratch. The authors state they do not yet know how far this generalizes.

    Ver el momento de apoyo · Párrafo 44

Pasajes clave4

Pasajes atribuidos con contexto para verificarlos. Abra el texto original para comprobar la fuente.

Generalization uncertainty

AI-driven kernel search reached near-expert performance on two kernels

Extracto original

On the two kernels we studied, AI-driven evolutionary kernel search grounded in structured cross-platform translation knowledge reached near-expert performance on Apple Silicon without a team of GPU experts starting from scratch. We do not yet know how far this generalizes
Contexto

, but the result is encouraging.

K-Search MLX performance results

Near-expert Apple Silicon performance via K-Search with MLX backend

Extracto original

We show that our approach reaches near-expert level performance on Apple Silicon with 0.97x speedup compared to the native MLX Attention kernel, and up to a 20x prefill speedup over the community mlx-lm implementation on the Mamba SSM kernel
Contexto

; we report the numbers, and how much of the gain comes from the translation layer, in the sections below. Although we focus on MLX kernels for Apple Silicon, the method is not specific to MLX and applies to any ecosystem where CUDA expertise is transferable.

Context quality over model capability

Bottleneck is context quality, not LLM coding ability

Extracto original

For us the main takeaway is that the bottleneck was not the LLM’s ability to write Metal code, but the quality of the context and constraints we gave it.
Contexto

Our CUDA translation layer converts existing NVIDIA kernel expertise into actionable guidance for Apple Silicon, and lets K-Search’s evolutionary search do the rest.

LLM-guided kernel porting

Naive LLM porting fails without hardware-aware constraints

Extracto original

Simply handing an LLM a CUDA kernel and asking it to port it is not enough: without deep hardware context, it produces code that is syntactically valid but architecturally wrong
Contexto

However, the more interesting challenge was not simply running K-Search on MLX. The key insight is that expert CUDA kernels encode decades of optimization knowledge that is transferable to Apple GPU if you can bridge the conceptual gap. (wrong tile sizes, invalid primitives, mismatched memory assumptions).

Fuente y metodología

Estas perspectivas enlazan a sus fuentes originales. Las paráfrasis están identificadas y no son citas textuales.

Abrir transcripción o material de origen (se abre en una pestaña nueva)Reportar un problema