公开观点

Kyle Wiggers

1 份资料 · 3 条观点 · 3 个话题

内容更新于:

Kyle Wiggers 关于MoE 扩展效率、数值精度优化、训练基础设施性能的观点。 按话题阅读 3 条观点,核对 1 个来源中的证据。

探索知识关联

按话题查看观点

按来源发布日期整理的个人观点,仅反映这些材料中的表达,不代表其全部立场。

译文仅辅助阅读;核查观点请以原始摘录为准。

MoE 扩展效率

查看话题
MoE 扩展效率

扩专家池,稳吞吐量

该文章报告称,将专家池从 8 个增加到 128 个,同时每个 token 仅选择四个专家,使每个 token 的激活参数大致保持在约 32 亿;总参数容量从 46 亿增长到 470 亿,而训练吞吐量下降不到 5%。

支持这项说法

发布 Olmo-core 3:面向大型 MoE 模型的开放、可扩展训练基础设施

在一项基准测试中,我们将专家池从 8 个增加到 128 个,同时仍为每个 token(即语言模型处理的小段文本单元)仅选择四个专家,使每个 token 的激活参数数量大致保持在约 32 亿。总参数容量从 46 亿增长到 470 亿,而训练吞吐量下降不到 5%。

原始摘录
In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.
上下文

Olmo-core 3 正是为了弥合这一差距而构建的。

原始上下文

Olmo-core 3 is built to close that gap.

训练基础设施性能

查看话题
训练基础设施性能

DDP MoE 堆栈吞吐量提升 2.7 倍

文章报告:在八块 NVIDIA B300 GPU 上,Olmo-core 3 新的分布式数据并行(DDP)MoE 训练堆栈达到每 GPU 每秒 52,000 个 token,而此前基于 FSDP 的实现为每 GPU 每秒 19,400 个 token,吞吐量提升约 2.7 倍。

支持这项说法

发布 Olmo-core 3:面向大型 MoE 模型的开放、可扩展训练基础设施

在八块 NVIDIA B300 GPU 上进行的初步测试中,一个 470 亿参数的 MoE 使用新堆栈每块 GPU 每秒处理 52,000 个 token,而我们早期的实现为 19,400 个——吞吐量约为 2.7 倍。

原始摘录
In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.
上下文

NVIDIA 的 Megatron-Core 是训练大型 MoE 的成熟方案。Olmo-core 3 为 Olmo 背后的框架带来了集成的 MoE 训练堆栈,其重新设计相比我们早期基于 FSDP 的实现提升了吞吐量。

原始上下文

NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.

数值精度优化

查看话题
数值精度优化

MXFP8 在受控 B300 基准测试中提升了吞吐量与峰值活跃内存表现

该博文报告称:在四块 NVIDIA B300 GPU 上开展的受控基准测试中(专家负载均匀分布),在系统中性能关键部分启用 MXFP8 后,端到端训练吞吐量较 BF16 基线提升了约 21%,峰值活跃内存则从 103 GiB 降至 95 GiB;大部分增益来源于前馈计算及专家间数据传输,而非仅来自注意力机制。

支持这项说法

发布 Olmo-core 3:面向大型 MoE 模型的开放、可扩展训练基础设施

在系统中帮助最大的部分启用 MXFP8 后,训练吞吐量比我们用作基线的更高精度格式 BF16 高出约 21%,同时峰值活跃内存从 103 GiB 降至 95 GiB。大部分增益来自前馈计算以及在专家之间移动数据,而非仅来自注意力机制。

原始摘录
With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.
上下文

我们在四块 NVIDIA B300 GPU 上进行了一项受控基准测试,以衡量 MXFP8 对端到端训练吞吐量的影响,并将工作均匀分配到各个专家。

原始上下文

We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts.

按来源日期阅读1

按原始来源的发布日期排序;措辞不同不代表立场发生变化。