让知识彼此相连
知识图谱
沿着人物与判断,回到原始证据。
1 位人物 · 1 个来源 · 1 条观点
放回语境
训练基础设施性能
选择一条判断,追溯到它所在的原始对话。
DDP MoE 堆栈吞吐量提升 2.7 倍
等分区块用于阅读定位,不表示排名或权重。
当前 1–1 / 共 1 条判断 · 按来源日期由新到旧
当前判断
DDP MoE 堆栈吞吐量提升 2.7 倍
文章报告:在八块 NVIDIA B300 GPU 上,Olmo-core 3 新的分布式数据并行(DDP)MoE 训练堆栈达到每 GPU 每秒 52,000 个 token,而此前基于 FSDP 的实现为每 GPU 每秒 19,400 个 token,吞吐量提升约 2.7 倍。
这些是个人表达的观点,并非共识度量。原始资料保持其原始语言。
支持这项说法
在八块 NVIDIA B300 GPU 上进行的初步测试中,一个 470 亿参数的 MoE 使用新堆栈每块 GPU 每秒处理 52,000 个 token,而我们早期的实现为 19,400 个——吞吐量约为 2.7 倍。
原始摘录
In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.
上下文
NVIDIA 的 Megatron-Core 是训练大型 MoE 的成熟方案。Olmo-core 3 为 Olmo 背后的框架带来了集成的 MoE 训练堆栈,其重新设计相比我们早期基于 FSDP 的实现提升了吞吐量。
原始上下文
NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.
日期表示来源发表时间,不代表观点发生变化。 缺少审核合格译文的内容保留原文。