Naive LLM porting fails without hardware-aware constraints
Prompting an LLM to port CUDA kernels to MLX/Metal produces syntactically valid but architecturally incorrect code unless guided by deep hardware context and explicit constraints.
Éléments favorables
Extrait original
Simply handing an LLM a CUDA kernel and asking it to port it is not enough: without deep hardware context, it produces code that is syntactically valid but architecturally wrong
Contexte
However, the more interesting challenge was not simply running K-Search on MLX. The key insight is that expert CUDA kernels encode decades of optimization knowledge that is transferable to Apple GPU if you can bridge the conceptual gap. (wrong tile sizes, invalid primitives, mismatched memory assumptions).