MXFP8 在受控 B300 基准测试中提升了吞吐量与峰值活跃内存表现
该博文报告称:在四块 NVIDIA B300 GPU 上开展的受控基准测试中(专家负载均匀分布),在系统中性能关键部分启用 MXFP8 后,端到端训练吞吐量较 BF16 基线提升了约 21%,峰值活跃内存则从 103 GiB 降至 95 GiB;大部分增益来源于前馈计算及专家间数据传输,而非仅来自注意力机制。
支持这项说法
在系统中帮助最大的部分启用 MXFP8 后,训练吞吐量比我们用作基线的更高精度格式 BF16 高出约 21%,同时峰值活跃内存从 103 GiB 降至 95 GiB。大部分增益来自前馈计算以及在专家之间移动数据,而非仅来自注意力机制。
原始摘录
With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.
上下文
我们在四块 NVIDIA B300 GPU 上进行了一项受控基准测试,以衡量 MXFP8 对端到端训练吞吐量的影响,并将工作均匀分配到各个专家。
原始上下文
We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts.