缺乏硬件感知约束时,简单的 LLM 移植会失败
提示 LLM 将 CUDA 内核移植到 MLX/Metal 会生成语法有效但架构不正确的代码,除非由深入的硬件上下文和显式约束加以引导。
支持这项说法
从 CUDA 到 MLX:K-Search 如何将数十年的内核专业知识引入 Apple Silicon
仅将一个 CUDA 内核交给 LLM 并要求其移植是不够的:若缺乏深入的硬件上下文,它生成的代码虽语法正确,但在架构层面是错误的。
原始摘录
Simply handing an LLM a CUDA kernel and asking it to port it is not enough: without deep hardware context, it produces code that is syntactically valid but architecturally wrong
上下文
然而,更具挑战性的问题并不只是在 MLX 上运行 K-Search。关键洞见在于,专家级 CUDA 内核编码了数十年的优化知识,如果你能弥合概念上的鸿沟,这些知识可以迁移到 Apple GPU。(错误的分块大小、无效的原语、不匹配的内存假设)。
原始上下文
However, the more interesting challenge was not simply running K-Search on MLX. The key insight is that expert CUDA kernels encode decades of optimization knowledge that is transferable to Apple GPU if you can bridge the conceptual gap. (wrong tile sizes, invalid primitives, mismatched memory assumptions).