ÖFFENTLICHE AUSGEDRÜCKTE MEINUNGEN

Kyle Wiggers

1 Quellen · 3 Standpunkte · 3 Themen

Inhalt aktualisiert:

Kyle Wiggers zu MoE scaling efficiency, numerical precision optimization, training infrastructure performance. Entdecke 3 Standpunkte nach Thema, mit Belegen aus 1 Quelle.

Zusammenhänge erkunden

Standpunkte nach Thema

Zugeordnete Standpunkte nach Veröffentlichungsdatum der Quelle. Eine Momentaufnahme dieser Beiträge, keine abschließende Darstellung persönlicher Überzeugungen.

Übersetzungen dienen dem Leseverständnis; die Originalauszüge bleiben die Quellenevidence.

MoE scaling efficiency

Thema ansehen

Maintaining throughput while scaling expert count

The post reports that increasing the expert pool from 8 to 128 while selecting only four experts per token kept active parameters per token roughly fixed at ~3.2B; total parameter capacity grew from 4.6B to 47B with less than 5% drop in training throughput.

Stützende Belege

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Originalauszug

In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.
Kontext

Olmo-core 3 is built to close that gap.

training infrastructure performance

Thema ansehen

2.7× throughput gain with DDP-based MoE stack

The post reports that on eight NVIDIA B300 GPUs, Olmo-core 3’s new distributed data parallelism (DDP)-based MoE training stack achieved 52,000 tokens/sec/GPU versus 19,400 tokens/sec/GPU with the prior FSDP-based implementation — a ~2.7× throughput improvement.

Stützende Belege

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Originalauszug

In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.
Kontext

NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.

numerical precision optimization

Thema ansehen

MXFP8 improves throughput and peak active memory in a controlled B300 benchmark

The post reports that in a controlled benchmark on four NVIDIA B300 GPUs with uniform expert load, enabling MXFP8 in performance-critical parts increased end-to-end training throughput by ~21% over BF16 baseline and reduced peak active memory from 103 GiB to 95 GiB; most gains came from feed-forward computation and inter-expert data movement, not attention alone.

Stützende Belege

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Originalauszug

With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.
Kontext

We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts.

Aussagen nach Quelldatum1

Die Aussagen sind nach dem Veröffentlichungsdatum der Originalquelle geordnet; unterschiedliche Formulierungen belegen keinen Positionswechsel.