VERNETZTES DENKEN

Wissensatlas

Personen, Standpunkte und ihre Originalbelege erkunden.

1 Personen · 1 Quellen · 1 geäußerte Meinungen

IM KONTEXT

MoE scaling efficiency

Einen Standpunkt auswählen und das Originalgespräch nachverfolgen.

Maintaining throughput while scaling expert count

IM KONTEXTMoE scaling efficiency1 ausgedrückte Meinungen
2026-10-01

Gleich große Sektoren dienen der Orientierung, nicht der Rangfolge.

Standpunkte 1–1 von 1 · Neueste Quellen zuerst

1 / 1

Ausgewählter Standpunkt

Maintaining throughput while scaling expert count

The post reports that increasing the expert pool from 8 to 128 while selecting only four experts per token kept active parameters per token roughly fixed at ~3.2B; total parameter capacity grew from 4.6B to 47B with less than 5% drop in training throughput.

Dies sind individuelle Standpunkte, keine Messung einer Übereinstimmung. Das Quellenmaterial bleibt in der Originalsprache.

Stützende Belege

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Originalauszug

In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.
Kontext

Olmo-core 3 is built to close that gap.

Die Daten beziehen sich auf Quellen, nicht auf Änderungen der Überzeugung. Texte ohne geprüfte Übersetzung bleiben in ihrer Originalsprache.