CONNECTED THINKING

Knowledge atlas

Follow people, viewpoints and their original evidence.

1 people · 1 sources · 1 viewpoints

IN CONTEXT

numerical precision optimization

Choose a viewpoint. Follow it back to the conversation.

MXFP8 improves throughput and peak active memory in a controlled B300 benchmark

IN CONTEXTnumerical precision optimization1 viewpoints
2026-10-01

Equal sectors are reading positions, not rankings.

Showing 1–1 of 1 viewpoints · Newest sources first

1 / 1

Selected viewpoint

MXFP8 improves throughput and peak active memory in a controlled B300 benchmark

The post reports that in a controlled benchmark on four NVIDIA B300 GPUs with uniform expert load, enabling MXFP8 in performance-critical parts increased end-to-end training throughput by ~21% over BF16 baseline and reduced peak active memory from 103 GiB to 95 GiB; most gains came from feed-forward computation and inter-expert data movement, not attention alone.

These are individual perspectives, not a measure of consensus. Source material stays in its original language.

Supporting evidence

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Original excerpt

With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.
Context

We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts.

Publication dates describe the sources, not changes in belief.