CONNECTED THINKING
Knowledge atlas
Follow people, viewpoints and their original evidence.
1 people · 1 sources · 1 viewpoints
IN CONTEXT
numerical precision optimization
Choose a viewpoint. Follow it back to the conversation.
MXFP8 improves throughput and peak active memory in a controlled B300 benchmark
Equal sectors are reading positions, not rankings.
Showing 1–1 of 1 viewpoints · Newest sources first
Selected viewpoint
MXFP8 improves throughput and peak active memory in a controlled B300 benchmark
The post reports that in a controlled benchmark on four NVIDIA B300 GPUs with uniform expert load, enabling MXFP8 in performance-critical parts increased end-to-end training throughput by ~21% over BF16 baseline and reduced peak active memory from 103 GiB to 95 GiB; most gains came from feed-forward computation and inter-expert data movement, not attention alone.
These are individual perspectives, not a measure of consensus. Source material stays in its original language.
Supporting evidence
Original excerpt
With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.
Context
We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts.
Publication dates describe the sources, not changes in belief.