发布 Olmo-core 3:面向大型 MoE 模型的开放、可扩展训练基础设施
在系统中帮助最大的部分启用 MXFP8 后,训练吞吐量比我们用作基线的更高精度格式 BF16 高出约 21%,同时峰值活跃内存从 103 GiB 降至 95 GiB。大部分增益来自前馈计算以及在专家之间移动数据,而非仅来自注意力机制。
原始摘录
With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.
上下文
我们在四块 NVIDIA B300 GPU 上进行了一项受控基准测试,以衡量 MXFP8 对端到端训练吞吐量的影响,并将工作均匀分配到各个专家。
原始上下文
We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts.