IDEAS CONECTADAS
Atlas de conocimiento
Explora personas, perspectivas y sus fuentes originales.
1 personas · 1 fuentes · 1 opiniones expresadas
EN CONTEXTO
training infrastructure performance
Elige una perspectiva y vuelve a la conversación original.
2.7× throughput gain with DDP-based MoE stack
Los sectores iguales orientan la lectura, no indican una clasificación.
Perspectivas 1–1 de 1 · Fuentes más recientes primero
Perspectiva seleccionada
2.7× throughput gain with DDP-based MoE stack
The post reports that on eight NVIDIA B300 GPUs, Olmo-core 3’s new distributed data parallelism (DDP)-based MoE training stack achieved 52,000 tokens/sec/GPU versus 19,400 tokens/sec/GPU with the prior FSDP-based implementation — a ~2.7× throughput improvement.
Estas son perspectivas individuales, no una medida de consenso. El material fuente permanece en su idioma original.
Evidencia a favor
Extracto original
In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.
Contexto
NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.
Las fechas corresponden a las fuentes, no a cambios de opinión. Los textos sin traducción revisada se mantienen en su idioma original.