UN THÈME, DANS SON CONTEXTE

MoE scaling efficiency

Judgments in this source concerning MoE scaling efficiency. Explorez 2 points de vue avec des éléments tirés de 2 sources.

1 personnes · 2 sources · 2 opinions exprimées

Contenu mis à jour:

Explorer les liens ↗

Carte des points de vue

Explorer par personne. Sélectionnez deux ou trois personnes pour les comparer.

1 personnes · 2 sources · 2 opinions exprimées

Kyle Wiggers

Maintaining throughput while scaling expert count

The post reports that increasing the expert pool from 8 to 128 while selecting only four experts per token kept active parameters per token roughly fixed at ~3.2B; total parameter capacity grew from 4.6B to 47B with less than 5% drop in training throughput.

Éléments favorables

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Extrait original

In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.
Contexte

Olmo-core 3 is built to close that gap.

Partager un aperçuVérifier cette affirmation

Maintaining throughput while scaling expert count

Increasing the expert pool from 8 to 128 while selecting only four experts per token kept active parameters per token roughly fixed at ~3.2B; total parameter capacity grew from 4.6B to 47B with less than 5% drop in training throughput.

Éléments favorables

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs | Ai2

Extrait original

In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.
Contexte

Olmo-core 3 is built to close that gap.

Ces résultats reflètent les sources disponibles, sans constituer une vue exhaustive ou à jour.