Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs | Ai2

Ai2 Research ·

Olmo-core 3 is presented as an open, scalable training infrastructure for large mixture-of-experts (MoE) models. The material reports three technical findings: maintaining throughput while scaling expert count from 8 to 128 with fixed active parameters per token; achieving ~2.7× higher training throughput than a prior FSDP-based MoE stack on eight NVIDIA B300 GPUs; and observing ~21% higher throughput and reduced peak GPU memory when using MXFP8 precision versus BF16 in a controlled benchmark. Read 3 viewpoints with supporting evidence and source links.

Ai2

Understand this piece

3 key points

Synthesis

  1. Maintaining throughput while scaling expert count

    Increasing the expert pool from 8 to 128 while selecting only four experts per token kept active parameters per token roughly fixed at ~3.2B; total parameter capacity grew from 4.6B to 47B with less than 5% drop in training throughput.

    Supporting evidence 1

    Original excerpt

    In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.

    Ai2 · Paragraph 4

    Context

    Olmo-core 3 is built to close that gap.

    Read in source context →

    Continue exploring

    MoE scaling efficiency →
  2. 2.7× throughput gain over prior FSDP-based MoE stack

    On eight NVIDIA B300 GPUs, Olmo-core 3 achieved 52,000 tokens/sec/GPU for a 47B-parameter MoE, compared to 19,400 tokens/sec/GPU with the earlier FSDP-based implementation — a ~2.7× improvement in throughput.

    Supporting evidence 1

    Original excerpt

    In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.

    Ai2 · Paragraph 10

    Context

    NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.

    Read in source context →
  3. MXFP8 yields 21% higher throughput and lower memory vs BF16

    In a controlled benchmark on four NVIDIA B300 GPUs with uniform expert load, enabling MXFP8 across optimal system components increased end-to-end training throughput by ~21% versus BF16 baseline, while peak active GPU memory decreased from 103 GiB to 95 GiB.

    Supporting evidence 1

    Original excerpt

    With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB.

    Ai2 · Paragraph 20

    Context

    We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.

    Read in source context →

Key passages3

Attributed passages with the context to verify them. Open the original text to check the source.

training infrastructure performance

2.7× throughput gain over prior FSDP-based MoE stack

Original excerpt

In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.
Context

NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.

MoE scaling efficiency

Maintaining throughput while scaling expert count

Original excerpt

In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.
Context

Olmo-core 3 is built to close that gap.

numerical precision optimization

MXFP8 yields 21% higher throughput and lower memory vs BF16

Original excerpt

With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB.
Context

We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.

Source & methodology

These viewpoints are linked to their original sources. Paraphrases are labeled and are not verbatim quotes.

Open transcript or source material (opens in a new tab)Report an issue

Continue with this topic

More sources on topics discussed here. Shared topics do not imply agreement.