Olmo-core 3 发布:面向大型 MoE 模型的开放、可扩展训练基础设施 | Ai2

Ai2 Research ·

Olmo-core 3 被介绍为一种面向大型混合专家(MoE)模型的开放、可扩展训练基础设施。该材料报告了三项技术发现:在保持每个 token 激活参数固定的情况下,将专家数量从 8 扩展到 128 时仍能维持吞吐量;在八块 NVIDIA B300 GPU 上实现比此前基于 FSDP 的 MoE 技术栈高约 2.7 倍的训练吞吐量;以及在受控基准测试中,使用 MXFP8 精度相较于 BF16 观察到吞吐量提高约 21% 并降低了 GPU 峰值显存。 阅读 3 条观点,查看支持证据与原始来源。

Ai2

理解这篇

3 个要点

综合解读

  1. 扩展专家数量的同时基本维持吞吐量

    将专家池从 8 增加到 128,同时每个 token 仅选择四个专家,使每个 token 的激活参数大致保持在约 32 亿;总参数容量从 46 亿增长到 470 亿,而训练吞吐量下降不到 5%。

    支持这项说法 1

    在一项基准测试中,我们将专家池从 8 增加到 128,同时每个 token 仍只选择四个专家——即语言模型处理的小段文本单位——使每个 token 的激活参数数量大致保持在约 32 亿。总参数容量从 46 亿增长到 470 亿,而训练吞吐量下降不到 5%。

    Ai2 · 段落 4

    原始摘录
    In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.
    上下文

    Olmo-core 3 正是为了弥合这一差距而构建的。

    原始上下文

    Olmo-core 3 is built to close that gap.

    回到原文语境 →
  2. 相较此前基于 FSDP 的 MoE 技术栈,吞吐量提升约 2.7 倍

    在八块 NVIDIA B300 GPU 上,Olmo-core 3 对一个 470 亿参数的 MoE 实现了每块 GPU 每秒 52,000 个 token,而早期基于 FSDP 的实现为每块 GPU 每秒 19,400 个 token——吞吐量提升约 2.7 倍。

    支持这项说法 1

    在八块 NVIDIA B300 GPU 上的初步测试中,一个 470 亿参数的 MoE 使用新技术栈实现了每块 GPU 每秒 52,000 个 token,而我们早期的实现为 19,400 个——吞吐量约为后者的 2.7 倍。

    Ai2 · 段落 10

    原始摘录
    In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.
    上下文

    NVIDIA 的 Megatron-Core 是训练大型 MoE 的成熟方案。Olmo-core 3 为 Olmo 背后的框架带来了集成的 MoE 训练技术栈,并通过重新设计提升了相较我们早期基于 FSDP 的实现的吞吐量。

    原始上下文

    NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.

    回到原文语境 →
  3. 与 BF16 相比,MXFP8 带来 21% 的更高吞吐量和更低显存占用

    在四块 NVIDIA B300 GPU 上进行的受控基准测试中,且专家负载均匀分布的情况下,在最优系统组件上启用 MXFP8 使端到端训练吞吐量较 BF16 基线提高约 21%,同时峰值活跃 GPU 显存从 103 GiB 降至 95 GiB。

    支持这项说法 1

    在系统中帮助最大的部分启用 MXFP8 后,训练吞吐量比我们用作基线的更高精度格式 BF16 高出约 21%,同时峰值活跃显存从 103 GiB 降至 95 GiB。

    Ai2 · 段落 20

    原始摘录
    With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB.
    上下文

    我们在四块 NVIDIA B300 GPU 上的受控基准测试中测量了 MXFP8 对端到端训练吞吐量的影响,工作负载在各专家之间均匀分配。大部分收益来自前馈计算和专家之间的数据移动,而非仅来自注意力机制。

    原始上下文

    We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.

    回到原文语境 →

关键段落3

带明确归属与语境的原文片段。打开原始文本核查出处。

训练基础设施性能

相较此前基于 FSDP 的 MoE 技术栈,吞吐量提升约 2.7 倍

在八块 NVIDIA B300 GPU 上的初步测试中,一个 470 亿参数的 MoE 使用新技术栈实现了每块 GPU 每秒 52,000 个 token,而我们早期的实现为 19,400 个——吞吐量约为后者的 2.7 倍。

原始摘录
In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.
上下文

NVIDIA 的 Megatron-Core 是训练大型 MoE 的成熟方案。Olmo-core 3 为 Olmo 背后的框架带来了集成的 MoE 训练技术栈,并通过重新设计提升了相较我们早期基于 FSDP 的实现的吞吐量。

原始上下文

NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.

MoE 扩展效率

扩展专家数量的同时基本维持吞吐量

在一项基准测试中,我们将专家池从 8 增加到 128,同时每个 token 仍只选择四个专家——即语言模型处理的小段文本单位——使每个 token 的激活参数数量大致保持在约 32 亿。总参数容量从 46 亿增长到 470 亿,而训练吞吐量下降不到 5%。

原始摘录
In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.
上下文

Olmo-core 3 正是为了弥合这一差距而构建的。

原始上下文

Olmo-core 3 is built to close that gap.

数值精度优化

与 BF16 相比,MXFP8 带来 21% 的更高吞吐量和更低显存占用

在系统中帮助最大的部分启用 MXFP8 后,训练吞吐量比我们用作基线的更高精度格式 BF16 高出约 21%,同时峰值活跃显存从 103 GiB 降至 95 GiB。

原始摘录
With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB.
上下文

我们在四块 NVIDIA B300 GPU 上的受控基准测试中测量了 MXFP8 对端到端训练吞吐量的影响,工作负载在各专家之间均匀分配。大部分收益来自前馈计算和专家之间的数据移动,而非仅来自注意力机制。

原始上下文

We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.

来源与研究方法

这些观点均关联原始来源。转述已明确标注,不作为逐字原话展示。

打开转录或来源材料 (在新标签页中打开)报告问题

继续研究相关话题

以下资料涉及本文的话题;讨论相同话题不代表观点一致。