OPINIONES PÚBLICAS

Kyle Wiggers

1 fuentes · 3 opiniones · 3 temas

Contenido actualizado:

Kyle Wiggers sobre MoE scaling efficiency, numerical precision optimization, training infrastructure performance. Explora 3 puntos de vista por tema, con evidencias de 1 fuente.

Explorar conexiones

Perspectivas por tema

Opiniones atribuidas, ordenadas por fecha de publicación de la fuente. Una muestra de estas intervenciones, no una definición completa de las creencias de la persona.

Las traducciones son para facilitar la lectura; los extractos originales siguen siendo la source evidence.

MoE scaling efficiency

Ver este tema

Maintaining throughput while scaling expert count

The post reports that increasing the expert pool from 8 to 128 while selecting only four experts per token kept active parameters per token roughly fixed at ~3.2B; total parameter capacity grew from 4.6B to 47B with less than 5% drop in training throughput.

Evidencia a favor

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Extracto original

In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.
Contexto

Olmo-core 3 is built to close that gap.

Kyle Wiggers
Compartir información

training infrastructure performance

Ver este tema

2.7× throughput gain with DDP-based MoE stack

The post reports that on eight NVIDIA B300 GPUs, Olmo-core 3’s new distributed data parallelism (DDP)-based MoE training stack achieved 52,000 tokens/sec/GPU versus 19,400 tokens/sec/GPU with the prior FSDP-based implementation — a ~2.7× throughput improvement.

Evidencia a favor

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Extracto original

In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.
Contexto

NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.

Kyle Wiggers
Compartir información

numerical precision optimization

Ver este tema

MXFP8 improves throughput and peak active memory in a controlled B300 benchmark

The post reports that in a controlled benchmark on four NVIDIA B300 GPUs with uniform expert load, enabling MXFP8 in performance-critical parts increased end-to-end training throughput by ~21% over BF16 baseline and reduced peak active memory from 103 GiB to 95 GiB; most gains came from feed-forward computation and inter-expert data movement, not attention alone.

Evidencia a favor

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Extracto original

With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.
Contexto

We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts.

Kyle Wiggers
Compartir información

Declaraciones por fecha de la fuente1

Las declaraciones se ordenan por fecha de publicación de la fuente original; las diferencias de redacción no demuestran un cambio de postura.