DDP MoE 堆栈吞吐量提升 2.7 倍
在八块 NVIDIA B300 GPU 上进行的初步测试中,一个 470 亿参数的 MoE 使用新堆栈每块 GPU 每秒处理 52,000 个 token,而我们早期的实现为 19,400 个——吞吐量约为 2.7 倍。
原始摘录
In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.
上下文
NVIDIA 的 Megatron-Core 是训练大型 MoE 的成熟方案。Olmo-core 3 为 Olmo 背后的框架带来了集成的 MoE 训练堆栈,其重新设计相比我们早期基于 FSDP 的实现提升了吞吐量。
原始上下文
NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.