VERNETZTES DENKEN

Wissensatlas

Personen, Standpunkte und ihre Originalbelege erkunden.

1 Personen · 1 Quellen · 1 geäußerte Meinungen

IM KONTEXT

training infrastructure performance

Einen Standpunkt auswählen und das Originalgespräch nachverfolgen.

2.7× throughput gain with DDP-based MoE stack

IM KONTEXTtraining infrastructure performance1 ausgedrückte Meinungen
2026-10-01

Gleich große Sektoren dienen der Orientierung, nicht der Rangfolge.

Standpunkte 1–1 von 1 · Neueste Quellen zuerst

1 / 1

Ausgewählter Standpunkt

2.7× throughput gain with DDP-based MoE stack

The post reports that on eight NVIDIA B300 GPUs, Olmo-core 3’s new distributed data parallelism (DDP)-based MoE training stack achieved 52,000 tokens/sec/GPU versus 19,400 tokens/sec/GPU with the prior FSDP-based implementation — a ~2.7× throughput improvement.

Dies sind individuelle Standpunkte, keine Messung einer Übereinstimmung. Das Quellenmaterial bleibt in der Originalsprache.

Stützende Belege

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Originalauszug

In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.
Kontext

NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.

Die Daten beziehen sich auf Quellen, nicht auf Änderungen der Überzeugung. Texte ohne geprüfte Übersetzung bleiben in ihrer Originalsprache.