瓶颈在于上下文质量,而非 LLM 编码能力
针对新硬件的 AI 驱动内核优化的主要瓶颈并非 LLM 生成 Metal 代码的能力,而是用于引导搜索的结构化跨平台翻译知识的质量。
支持这项说法
从 CUDA 到 MLX:K-Search 如何将数十年的内核专业知识引入 Apple Silicon
对我们而言,主要结论在于:瓶颈并非大语言模型(LLM)编写 Metal 代码的能力,而在于我们为其提供的上下文与约束条件的质量。
原始摘录
For us the main takeaway is that the bottleneck was not the LLM’s ability to write Metal code, but the quality of the context and constraints we gave it.
上下文
我们的 CUDA 转换层将现有的 NVIDIA 内核专业知识转化为针对 Apple Silicon 的可操作指导,并由 K-Search 的进化式搜索完成其余工作。
原始上下文
Our CUDA translation layer converts existing NVIDIA kernel expertise into actionable guidance for Apple Silicon, and lets K-Search’s evolutionary search do the rest.