VERNETZTES DENKEN

Wissensatlas

Personen, Standpunkte und ihre Originalbelege erkunden.

1 Personen · 1 Quellen · 1 geäußerte Meinungen

IM KONTEXT

numerical precision optimization

Einen Standpunkt auswählen und das Originalgespräch nachverfolgen.

MXFP8 improves throughput and peak active memory in a controlled B300 benchmark

IM KONTEXTnumerical precision optimization1 ausgedrückte Meinungen
2026-10-01

Gleich große Sektoren dienen der Orientierung, nicht der Rangfolge.

Standpunkte 1–1 von 1 · Neueste Quellen zuerst

1 / 1

Ausgewählter Standpunkt

MXFP8 improves throughput and peak active memory in a controlled B300 benchmark

The post reports that in a controlled benchmark on four NVIDIA B300 GPUs with uniform expert load, enabling MXFP8 in performance-critical parts increased end-to-end training throughput by ~21% over BF16 baseline and reduced peak active memory from 103 GiB to 95 GiB; most gains came from feed-forward computation and inter-expert data movement, not attention alone.

Dies sind individuelle Standpunkte, keine Messung einer Übereinstimmung. Das Quellenmaterial bleibt in der Originalsprache.

Stützende Belege

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Originalauszug

With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.
Kontext

We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts.

Die Daten beziehen sich auf Quellen, nicht auf Änderungen der Überzeugung. Texte ohne geprüfte Übersetzung bleiben in ihrer Originalsprache.