VERNETZTES DENKEN
Wissensatlas
Personen, Standpunkte und ihre Originalbelege erkunden.
1 Personen · 1 Quellen · 1 geäußerte Meinungen
IM KONTEXT
numerical precision optimization
Einen Standpunkt auswählen und das Originalgespräch nachverfolgen.
MXFP8 improves throughput and peak active memory in a controlled B300 benchmark
Gleich große Sektoren dienen der Orientierung, nicht der Rangfolge.
Standpunkte 1–1 von 1 · Neueste Quellen zuerst
Ausgewählter Standpunkt
MXFP8 improves throughput and peak active memory in a controlled B300 benchmark
The post reports that in a controlled benchmark on four NVIDIA B300 GPUs with uniform expert load, enabling MXFP8 in performance-critical parts increased end-to-end training throughput by ~21% over BF16 baseline and reduced peak active memory from 103 GiB to 95 GiB; most gains came from feed-forward computation and inter-expert data movement, not attention alone.
Dies sind individuelle Standpunkte, keine Messung einer Übereinstimmung. Das Quellenmaterial bleibt in der Originalsprache.
Stützende Belege
Originalauszug
With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.
Kontext
We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts.
Die Daten beziehen sich auf Quellen, nicht auf Änderungen der Überzeugung. Texte ohne geprüfte Übersetzung bleiben in ihrer Originalsprache.