IDEAS CONECTADAS

Atlas de conocimiento

Explora personas, perspectivas y sus fuentes originales.

1 personas · 1 fuentes · 1 opiniones expresadas

EN CONTEXTO

MoE scaling efficiency

Elige una perspectiva y vuelve a la conversación original.

Maintaining throughput while scaling expert count

EN CONTEXTOMoE scaling efficiency1 opiniones expresadas
2026-10-01

Los sectores iguales orientan la lectura, no indican una clasificación.

Perspectivas 1–1 de 1 · Fuentes más recientes primero

1 / 1

Perspectiva seleccionada

Maintaining throughput while scaling expert count

The post reports that increasing the expert pool from 8 to 128 while selecting only four experts per token kept active parameters per token roughly fixed at ~3.2B; total parameter capacity grew from 4.6B to 47B with less than 5% drop in training throughput.

Estas son perspectivas individuales, no una medida de consenso. El material fuente permanece en su idioma original.

Evidencia a favor

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Extracto original

In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.
Contexto

Olmo-core 3 is built to close that gap.

Las fechas corresponden a las fuentes, no a cambios de opinión. Los textos sin traducción revisada se mantienen en su idioma original.