UN TEMA, EN CONTEXTO

K-Search MLX performance results

Judgments in this source concerning K-Search MLX performance results. Explora 1 punto de vista con evidencias de 1 fuente.

0 personas · 1 fuentes · 1 opiniones expresadas

Contenido actualizado:

Explorar conexiones ↗

Mapa de perspectivas

0 personas · 1 fuentes · 1 opiniones expresadas

Near-expert Apple Silicon performance via K-Search with MLX backend

K-Search with a novel CUDA-to-MLX translation layer achieved 0.97× speedup versus the native MLX Attention kernel and up to 20× prefill speedup over mlx-lm on the Mamba SSM kernel on Apple Silicon. The authors note the method applies to any ecosystem where CUDA expertise is transferable.

Evidencia a favor

From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon

Extracto original

We show that our approach reaches near-expert level performance on Apple Silicon with 0.97x speedup compared to the native MLX Attention kernel, and up to a 20x prefill speedup over the community mlx-lm implementation on the Mamba SSM kernel
Contexto

; we report the numbers, and how much of the gain comes from the translation layer, in the sections below. Although we focus on MLX kernels for Apple Silicon, the method is not specific to MLX and applies to any ecosystem where CUDA expertise is transferable.

Estos resultados reflejan las fuentes disponibles, no una visión exhaustiva ni necesariamente actual.