Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
Originalauszug
In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.
Kontext
Olmo-core 3 is built to close that gap.