EIN THEMA, IM KONTEXT

numerical precision optimization

Judgments in this source concerning numerical precision optimization. Entdecke 2 Standpunkte mit Belegen aus 2 Quellen.

1 Personen · 2 Quellen · 2 geäußerte Meinungen

Inhalt aktualisiert:

Zusammenhänge erkunden ↗

Perspektiven im Überblick

Erkunden Sie nach Person. Wählen Sie zwei oder drei zum Vergleich aus.

1 Personen · 2 Quellen · 2 geäußerte Meinungen

Kyle Wiggers

MXFP8 improves throughput and peak active memory in a controlled B300 benchmark

The post reports that in a controlled benchmark on four NVIDIA B300 GPUs with uniform expert load, enabling MXFP8 in performance-critical parts increased end-to-end training throughput by ~21% over BF16 baseline and reduced peak active memory from 103 GiB to 95 GiB; most gains came from feed-forward computation and inter-expert data movement, not attention alone.

Stützende Belege

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Originalauszug

With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.
Kontext

We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts.

Erkenntnisse teilenDiese Aussage überprüfen

MXFP8 yields 21% higher throughput and lower memory vs BF16

In a controlled benchmark on four NVIDIA B300 GPUs with uniform expert load, enabling MXFP8 across optimal system components increased end-to-end training throughput by ~21% versus BF16 baseline, while peak active GPU memory decreased from 103 GiB to 95 GiB.

Stützende Belege

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs | Ai2

Originalauszug

With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB.
Kontext

We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.

Diese Ergebnisse spiegeln die verfügbaren Quellen wider, nicht ein vollständiges oder aktuelles Bild.