扩专家池,稳吞吐量
该文章报告称,将专家池从 8 个增加到 128 个,同时每个 token 仅选择四个专家,使每个 token 的激活参数大致保持在约 32 亿;总参数容量从 46 亿增长到 470 亿,而训练吞吐量下降不到 5%。
支持这项说法
在一项基准测试中,我们将专家池从 8 个增加到 128 个,同时仍为每个 token(即语言模型处理的小段文本单元)仅选择四个专家,使每个 token 的激活参数数量大致保持在约 32 亿。总参数容量从 46 亿增长到 470 亿,而训练吞吐量下降不到 5%。
原始摘录
In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.
上下文
Olmo-core 3 正是为了弥合这一差距而构建的。
原始上下文
Olmo-core 3 is built to close that gap.