From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon

Berkeley BAIR Blog ·

A technical report on K-Search, a method for porting CUDA kernels to Apple Silicon using MLX and Metal. It focuses on hardware-aware constraints in LLM-guided translation, identifies context quality—not model capability—as the key bottleneck, reports performance relative to native MLX kernels, and notes observed success on two kernels (attention and Mamba SSM), while acknowledging uncertainty about broader generalization. Read 4 viewpoints with supporting evidence and source links.

Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez

Understand this piece

4 key points

Synthesis

  1. Naive LLM porting fails without hardware-aware constraints

    Prompting an LLM to port CUDA kernels to MLX/Metal produces syntactically valid but architecturally incorrect code unless guided by deep hardware context and explicit constraints.

    Supporting evidence 1

    Original excerpt

    Simply handing an LLM a CUDA kernel and asking it to port it is not enough: without deep hardware context, it produces code that is syntactically valid but architecturally wrong

    Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez · Paragraph 26

    Context

    However, the more interesting challenge was not simply running K-Search on MLX. The key insight is that expert CUDA kernels encode decades of optimization knowledge that is transferable to Apple GPU if you can bridge the conceptual gap. (wrong tile sizes, invalid primitives, mismatched memory assumptions).

    Read in source context →
  2. Bottleneck is context quality, not LLM coding ability

    The main bottleneck in AI-driven kernel optimization for new hardware is not the LLM’s ability to generate Metal code, but the quality of the structured, cross-platform translation knowledge used to guide the search.

    Supporting evidence 1

    Original excerpt

    For us the main takeaway is that the bottleneck was not the LLM’s ability to write Metal code, but the quality of the context and constraints we gave it.

    Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez · Paragraph 45

    Context

    Our CUDA translation layer converts existing NVIDIA kernel expertise into actionable guidance for Apple Silicon, and lets K-Search’s evolutionary search do the rest.

    Read in source context →
  3. Near-expert Apple Silicon performance via K-Search with MLX backend

    K-Search with a novel CUDA-to-MLX translation layer achieved 0.97× speedup versus the native MLX Attention kernel and up to 20× prefill speedup over mlx-lm on the Mamba SSM kernel on Apple Silicon. The authors note the method applies to any ecosystem where CUDA expertise is transferable.

    Supporting evidence 1

    Original excerpt

    We show that our approach reaches near-expert level performance on Apple Silicon with 0.97x speedup compared to the native MLX Attention kernel, and up to a 20x prefill speedup over the community mlx-lm implementation on the Mamba SSM kernel

    Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez · Paragraph 5

    Context

    ; we report the numbers, and how much of the gain comes from the translation layer, in the sections below. Although we focus on MLX kernels for Apple Silicon, the method is not specific to MLX and applies to any ecosystem where CUDA expertise is transferable.

    Read in source context →
  4. AI-driven kernel search reached near-expert performance on two kernels

    On the two kernels studied—attention and Mamba SSM—AI-driven evolutionary kernel search, grounded in structured cross-platform translation, reached near-expert performance on Apple Silicon without requiring GPU experts to start from scratch. The authors state they do not yet know how far this generalizes.

    Supporting evidence 1

    Original excerpt

    On the two kernels we studied, AI-driven evolutionary kernel search grounded in structured cross-platform translation knowledge reached near-expert performance on Apple Silicon without a team of GPU experts starting from scratch. We do not yet know how far this generalizes

    Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez · Paragraph 44

    Context

    , but the result is encouraging.

    Read in source context →

Key passages4

Attributed passages with the context to verify them. Open the original text to check the source.

Generalization uncertainty

AI-driven kernel search reached near-expert performance on two kernels

Original excerpt

On the two kernels we studied, AI-driven evolutionary kernel search grounded in structured cross-platform translation knowledge reached near-expert performance on Apple Silicon without a team of GPU experts starting from scratch. We do not yet know how far this generalizes
Context

, but the result is encouraging.

K-Search MLX performance results

Near-expert Apple Silicon performance via K-Search with MLX backend

Original excerpt

We show that our approach reaches near-expert level performance on Apple Silicon with 0.97x speedup compared to the native MLX Attention kernel, and up to a 20x prefill speedup over the community mlx-lm implementation on the Mamba SSM kernel
Context

; we report the numbers, and how much of the gain comes from the translation layer, in the sections below. Although we focus on MLX kernels for Apple Silicon, the method is not specific to MLX and applies to any ecosystem where CUDA expertise is transferable.

Context quality over model capability

Bottleneck is context quality, not LLM coding ability

Original excerpt

For us the main takeaway is that the bottleneck was not the LLM’s ability to write Metal code, but the quality of the context and constraints we gave it.
Context

Our CUDA translation layer converts existing NVIDIA kernel expertise into actionable guidance for Apple Silicon, and lets K-Search’s evolutionary search do the rest.

LLM-guided kernel porting

Naive LLM porting fails without hardware-aware constraints

Original excerpt

Simply handing an LLM a CUDA kernel and asking it to port it is not enough: without deep hardware context, it produces code that is syntactically valid but architecturally wrong
Context

However, the more interesting challenge was not simply running K-Search on MLX. The key insight is that expert CUDA kernels encode decades of optimization knowledge that is transferable to Apple GPU if you can bridge the conceptual gap. (wrong tile sizes, invalid primitives, mismatched memory assumptions).

Mentioned here

All mentioned things

MLX

Mention only

The authors report performance results for their method on MLX kernels (Attention and Mamba SSM) on Apple Silicon, treating MLX as the target ecosystem—not as a tool under critique or endorsement.

Read supporting evidence · Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez

Source & methodology

These viewpoints are linked to their original sources. Paraphrases are labeled and are not verbatim quotes.

Open transcript or source material (opens in a new tab)Report an issue