UN THÈME, DANS SON CONTEXTE

training infrastructure performance

Judgments in this source concerning training infrastructure performance. Explorez 2 points de vue avec des éléments tirés de 2 sources.

1 personnes · 2 sources · 2 opinions exprimées

Contenu mis à jour:

Explorer les liens ↗

Carte des points de vue

Explorer par personne. Sélectionnez deux ou trois personnes pour les comparer.

1 personnes · 2 sources · 2 opinions exprimées

Kyle Wiggers

2.7× throughput gain with DDP-based MoE stack

The post reports that on eight NVIDIA B300 GPUs, Olmo-core 3’s new distributed data parallelism (DDP)-based MoE training stack achieved 52,000 tokens/sec/GPU versus 19,400 tokens/sec/GPU with the prior FSDP-based implementation — a ~2.7× throughput improvement.

Éléments favorables

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Extrait original

In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.
Contexte

NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.

Partager un aperçuVérifier cette affirmation

2.7× throughput gain over prior FSDP-based MoE stack

On eight NVIDIA B300 GPUs, Olmo-core 3 achieved 52,000 tokens/sec/GPU for a 47B-parameter MoE, compared to 19,400 tokens/sec/GPU with the earlier FSDP-based implementation — a ~2.7× improvement in throughput.

Éléments favorables

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs | Ai2

Extrait original

In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.
Contexte

NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.

Ces résultats reflètent les sources disponibles, sans constituer une vue exhaustive ou à jour.