CONNECTED THINKING
Knowledge atlas
Follow people, viewpoints and their original evidence.
1 people · 1 sources · 1 viewpoints
IN CONTEXT
MoE scaling efficiency
Choose a viewpoint. Follow it back to the conversation.
Maintaining throughput while scaling expert count
Equal sectors are reading positions, not rankings.
Showing 1–1 of 1 viewpoints · Newest sources first
Selected viewpoint
Maintaining throughput while scaling expert count
The post reports that increasing the expert pool from 8 to 128 while selecting only four experts per token kept active parameters per token roughly fixed at ~3.2B; total parameter capacity grew from 4.6B to 47B with less than 5% drop in training throughput.
These are individual perspectives, not a measure of consensus. Source material stays in its original language.
Supporting evidence
Original excerpt
In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.
Context
Olmo-core 3 is built to close that gap.
Publication dates describe the sources, not changes in belief.