IDEAS CONECTADAS
Atlas de conocimiento
Explora personas, perspectivas y sus fuentes originales.
1 personas · 1 fuentes · 1 opiniones expresadas
EN CONTEXTO
numerical precision optimization
Elige una perspectiva y vuelve a la conversación original.
MXFP8 improves throughput and peak active memory in a controlled B300 benchmark
Los sectores iguales orientan la lectura, no indican una clasificación.
Perspectivas 1–1 de 1 · Fuentes más recientes primero
Perspectiva seleccionada
MXFP8 improves throughput and peak active memory in a controlled B300 benchmark
The post reports that in a controlled benchmark on four NVIDIA B300 GPUs with uniform expert load, enabling MXFP8 in performance-critical parts increased end-to-end training throughput by ~21% over BF16 baseline and reduced peak active memory from 103 GiB to 95 GiB; most gains came from feed-forward computation and inter-expert data movement, not attention alone.
Estas son perspectivas individuales, no una medida de consenso. El material fuente permanece en su idioma original.
Evidencia a favor
Extracto original
With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.
Contexto
We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts.
Las fechas corresponden a las fuentes, no a cambios de opinión. Los textos sin traducción revisada se mantienen en su idioma original.