AI-driven kernel search reached near-expert performance on two kernels
On the two kernels studied—attention and Mamba SSM—AI-driven evolutionary kernel search, grounded in structured cross-platform translation, reached near-expert performance on Apple Silicon without requiring GPU experts to start from scratch. The authors state they do not yet know how far this generalizes.
Éléments favorables
From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon
Extrait original
On the two kernels we studied, AI-driven evolutionary kernel search grounded in structured cross-platform translation knowledge reached near-expert performance on Apple Silicon without a team of GPU experts starting from scratch. We do not yet know how far this generalizes
Contexte
, but the result is encouraging.