相较此前基于 FSDP 的 MoE 技术栈,吞吐量提升约 2.7 倍
在八块 NVIDIA B300 GPU 上的初步测试中,一个 470 亿参数的 MoE 使用新技术栈实现了每块 GPU 每秒 52,000 个 token,而我们早期的实现为 19,400 个——吞吐量约为后者的 2.7 倍。
原始摘录
In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.
上下文
NVIDIA 的 Megatron-Core 是训练大型 MoE 的成熟方案。Olmo-core 3 为 Olmo 背后的框架带来了集成的 MoE 训练技术栈,并通过重新设计提升了相较我们早期基于 FSDP 的实现的吞吐量。
原始上下文
NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.