发布 Olmo-core 3:面向大型 MoE 模型的开放、可扩展训练基础设施

Hugging Face Blog ·

Kyle Wiggers 在 Hugging Face 上的文章报道了 Ai2 自有的 Olmo-core 3 基准测试。文中描述了在激活参数大致固定的情况下扩展专家池、与早期堆栈进行的初步八 GPU 吞吐量对比,以及与 BF16 的受控 MXFP8 对比。这些均为该文章中报告的基准测试结果,而非独立测量数据。 阅读 3 条观点,查看支持证据与原始来源。

理解这篇

3 个要点

综合解读

  1. 扩专家池,稳吞吐量

    该文章报告称,将专家池从 8 个增加到 128 个,同时每个 token 仅选择四个专家,使每个 token 的激活参数大致保持在约 32 亿;总参数容量从 46 亿增长到 470 亿,而训练吞吐量下降不到 5%。

    支持这项说法 1

    在一项基准测试中,我们将专家池从 8 个增加到 128 个,同时仍为每个 token(即语言模型处理的小段文本单元)仅选择四个专家,使每个 token 的激活参数数量大致保持在约 32 亿。总参数容量从 46 亿增长到 470 亿,而训练吞吐量下降不到 5%。

    Kyle Wiggers · 段落 5

    原始摘录
    In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.
    上下文

    Olmo-core 3 正是为了弥合这一差距而构建的。

    原始上下文

    Olmo-core 3 is built to close that gap.

    回到原文语境 →
  2. DDP MoE 堆栈吞吐量提升 2.7 倍

    文章报告:在八块 NVIDIA B300 GPU 上,Olmo-core 3 新的分布式数据并行(DDP)MoE 训练堆栈达到每 GPU 每秒 52,000 个 token,而此前基于 FSDP 的实现为每 GPU 每秒 19,400 个 token,吞吐量提升约 2.7 倍。

    支持这项说法 1

    在八块 NVIDIA B300 GPU 上进行的初步测试中,一个 470 亿参数的 MoE 使用新堆栈每块 GPU 每秒处理 52,000 个 token,而我们早期的实现为 19,400 个——吞吐量约为 2.7 倍。

    Kyle Wiggers · 段落 11

    原始摘录
    In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.
    上下文

    NVIDIA 的 Megatron-Core 是训练大型 MoE 的成熟方案。Olmo-core 3 为 Olmo 背后的框架带来了集成的 MoE 训练堆栈,其重新设计相比我们早期基于 FSDP 的实现提升了吞吐量。

    原始上下文

    NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.

    回到原文语境 →
  3. MXFP8 在受控 B300 基准测试中提升了吞吐量与峰值活跃内存表现

    该博文报告称:在四块 NVIDIA B300 GPU 上开展的受控基准测试中(专家负载均匀分布),在系统中性能关键部分启用 MXFP8 后,端到端训练吞吐量较 BF16 基线提升了约 21%,峰值活跃内存则从 103 GiB 降至 95 GiB;大部分增益来源于前馈计算及专家间数据传输,而非仅来自注意力机制。

    支持这项说法 1

    在系统中帮助最大的部分启用 MXFP8 后,训练吞吐量比我们用作基线的更高精度格式 BF16 高出约 21%,同时峰值活跃内存从 103 GiB 降至 95 GiB。大部分增益来自前馈计算以及在专家之间移动数据,而非仅来自注意力机制。

    Kyle Wiggers · 段落 21

    原始摘录
    With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.
    上下文

    我们在四块 NVIDIA B300 GPU 上进行了一项受控基准测试,以衡量 MXFP8 对端到端训练吞吐量的影响,并将工作均匀分配到各个专家。

    原始上下文

    We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts.

    回到原文语境 →

关键段落3

带明确归属与语境的原文片段。打开原始文本核查出处。

训练基础设施性能

DDP MoE 堆栈吞吐量提升 2.7 倍

在八块 NVIDIA B300 GPU 上进行的初步测试中,一个 470 亿参数的 MoE 使用新堆栈每块 GPU 每秒处理 52,000 个 token,而我们早期的实现为 19,400 个——吞吐量约为 2.7 倍。

原始摘录
In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.
上下文

NVIDIA 的 Megatron-Core 是训练大型 MoE 的成熟方案。Olmo-core 3 为 Olmo 背后的框架带来了集成的 MoE 训练堆栈,其重新设计相比我们早期基于 FSDP 的实现提升了吞吐量。

原始上下文

NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.

MoE 扩展效率

扩专家池,稳吞吐量

在一项基准测试中,我们将专家池从 8 个增加到 128 个,同时仍为每个 token(即语言模型处理的小段文本单元)仅选择四个专家,使每个 token 的激活参数数量大致保持在约 32 亿。总参数容量从 46 亿增长到 470 亿,而训练吞吐量下降不到 5%。

原始摘录
In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.
上下文

Olmo-core 3 正是为了弥合这一差距而构建的。

原始上下文

Olmo-core 3 is built to close that gap.

数值精度优化

MXFP8 在受控 B300 基准测试中提升了吞吐量与峰值活跃内存表现

在系统中帮助最大的部分启用 MXFP8 后,训练吞吐量比我们用作基线的更高精度格式 BF16 高出约 21%,同时峰值活跃内存从 103 GiB 降至 95 GiB。大部分增益来自前馈计算以及在专家之间移动数据,而非仅来自注意力机制。

原始摘录
With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.
上下文

我们在四块 NVIDIA B300 GPU 上进行了一项受控基准测试,以衡量 MXFP8 对端到端训练吞吐量的影响,并将工作均匀分配到各个专家。

原始上下文

We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts.

来源与研究方法

这些观点均关联原始来源。转述已明确标注,不作为逐字原话展示。

打开转录或来源材料 (在新标签页中打开)报告问题

继续了解这些人物的观点

继续研究相关话题

以下资料涉及本文的话题;讨论相同话题不代表观点一致。