From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon

Berkeley BAIR Blog ·

A technical report on K-Search, a method for porting CUDA kernels to Apple Silicon using MLX and Metal. It focuses on hardware-aware constraints in LLM-guided translation, identifies context quality—not model capability—as the key bottleneck, reports performance relative to native MLX kernels, and notes observed success on two kernels (attention and Mamba SSM), while acknowledging uncertainty about broader generalization. Lisez 4 points de vue avec leurs éléments à l’appui et les liens vers les sources.

Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez

En un coup d’œil

  • Naive LLM porting fails without hardware-aware constraints

    Prompting an LLM to port CUDA kernels to MLX/Metal produces syntactically valid but architecturally incorrect code unless guided by deep hardware context and explicit constraints.

    Lire le moment probant · Paragraphe 26
  • Bottleneck is context quality, not LLM coding ability

    The main bottleneck in AI-driven kernel optimization for new hardware is not the LLM’s ability to generate Metal code, but the quality of the structured, cross-platform translation knowledge used to guide the search.

    Lire le moment probant · Paragraphe 45
  • Near-expert Apple Silicon performance via K-Search with MLX backend

    K-Search with a novel CUDA-to-MLX translation layer achieved 0.97× speedup versus the native MLX Attention kernel and up to 20× prefill speedup over mlx-lm on the Mamba SSM kernel on Apple Silicon. The authors note the method applies to any ecosystem where CUDA expertise is transferable.

    Lire le moment probant · Paragraphe 5
  • AI-driven kernel search reached near-expert performance on two kernels

    On the two kernels studied—attention and Mamba SSM—AI-driven evolutionary kernel search, grounded in structured cross-platform translation, reached near-expert performance on Apple Silicon without requiring GPU experts to start from scratch. The authors state they do not yet know how far this generalizes.

    Lire le moment probant · Paragraphe 44

Passages clés4

Passages attribués et accompagnés du contexte nécessaire à leur vérification. Ouvrez le texte original pour vérifier la source.

Generalization uncertainty

AI-driven kernel search reached near-expert performance on two kernels

Extrait original

On the two kernels we studied, AI-driven evolutionary kernel search grounded in structured cross-platform translation knowledge reached near-expert performance on Apple Silicon without a team of GPU experts starting from scratch. We do not yet know how far this generalizes
Contexte

, but the result is encouraging.

K-Search MLX performance results

Near-expert Apple Silicon performance via K-Search with MLX backend

Extrait original

We show that our approach reaches near-expert level performance on Apple Silicon with 0.97x speedup compared to the native MLX Attention kernel, and up to a 20x prefill speedup over the community mlx-lm implementation on the Mamba SSM kernel
Contexte

; we report the numbers, and how much of the gain comes from the translation layer, in the sections below. Although we focus on MLX kernels for Apple Silicon, the method is not specific to MLX and applies to any ecosystem where CUDA expertise is transferable.

Context quality over model capability

Bottleneck is context quality, not LLM coding ability

Extrait original

For us the main takeaway is that the bottleneck was not the LLM’s ability to write Metal code, but the quality of the context and constraints we gave it.
Contexte

Our CUDA translation layer converts existing NVIDIA kernel expertise into actionable guidance for Apple Silicon, and lets K-Search’s evolutionary search do the rest.

LLM-guided kernel porting

Naive LLM porting fails without hardware-aware constraints

Extrait original

Simply handing an LLM a CUDA kernel and asking it to port it is not enough: without deep hardware context, it produces code that is syntactically valid but architecturally wrong
Contexte

However, the more interesting challenge was not simply running K-Search on MLX. The key insight is that expert CUDA kernels encode decades of optimization knowledge that is transferable to Apple GPU if you can bridge the conceptual gap. (wrong tile sizes, invalid primitives, mismatched memory assumptions).

Source et méthodologie

Ces points de vue renvoient à leurs sources originales. Les reformulations sont signalées et ne sont pas des citations mot à mot.

Ouvrir la transcription ou les documents sources (s’ouvre dans un nouvel onglet)Signaler un problème