CONNECTED THINKING
Knowledge atlas
Follow people, viewpoints and their original evidence.
1 people · 1 sources · 1 viewpoints
IN CONTEXT
training infrastructure performance
Choose a viewpoint. Follow it back to the conversation.
2.7× throughput gain with DDP-based MoE stack
Equal sectors are reading positions, not rankings.
Showing 1–1 of 1 viewpoints · Newest sources first
Selected viewpoint
2.7× throughput gain with DDP-based MoE stack
The post reports that on eight NVIDIA B300 GPUs, Olmo-core 3’s new distributed data parallelism (DDP)-based MoE training stack achieved 52,000 tokens/sec/GPU versus 19,400 tokens/sec/GPU with the prior FSDP-based implementation — a ~2.7× throughput improvement.
These are individual perspectives, not a measure of consensus. Source material stays in its original language.
Supporting evidence
Original excerpt
In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.
Context
NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.
Publication dates describe the sources, not changes in belief.