Mixture of Experts from First Principles
Breaking the link
From the FLOPs note, a dense model costs $6N$ training FLOPs per token, where $N$ counts all parameters. Scaling capacity therefore scales cost one-for-one.
An MoE layer replaces the single feed-forward network with $E$ independent copies — the experts — plus a small router that picks which $k$ of them each token visits.
| Configuration | Experts $E$ | Top-$k$ | Total FFN params | Active per token | Sparsity |
|---|---|---|---|---|---|
| Dense | 1 | 1 | baseline | same | 100% |
| 8 large experts | 8 | 2 | 45.1B | 11.3B | 25% |
| 256 fine experts + 1 shared | 256 | 8 | 656.5B | 23.0B | 3.5% |
The second row buys 29 times the parameters for the FLOPs of a much smaller model. That is the entire pitch, and the rest of this note is about what it costs.
The router
The router is a single linear layer, $E$ outputs wide:
and the chosen experts' outputs are combined with their (renormalized) gate values as weights. It is a tiny amount of arithmetic guarding an enormous amount of parameter.
Load balancing
The standard remedy is an auxiliary loss that penalizes uneven routing. Let $f_i$ be the fraction of tokens dispatched to expert $i$ and $P_i$ the mean router probability assigned to it:
The $f_i$ term is a count and carries no gradient; $P_i$ does. Pushing down $P_i$ for over-subscribed experts is what nudges the router toward balance. The multiplier $E$ normalizes the loss so that a perfectly uniform distribution gives $\mathcal L_{\text{aux}}=1$ for any $E$, which makes $\alpha$ transferable across configurations.
$\alpha$ too small
Experts collapse onto a few favourites. Effective capacity drops to a fraction of what you paid for, and the extra parameters are dead weight.
$\alpha$ too large
Routing becomes near-uniform, which means near-random. Experts cannot specialize, and the model behaves like a noisier dense one.
Because the auxiliary loss perturbs the objective you actually care about, recent systems prefer a bias-based alternative: keep a per-expert bias added to the routing scores only for the top-$k$ selection, and adjust it up or down between steps according to observed load. Balance is enforced without touching the gradient of the language-modelling loss at all.
Capacity and dropped tokens
Balance matters for a second, purely mechanical reason. Expert buffers must be allocated with a fixed shape before routing is known, so each expert gets room for a bounded number of tokens:
where $T$ is the tokens in the batch and $kT/E$ is the average load. A capacity factor of $1.25$ gives 25% headroom.
Overflow
Tokens beyond an expert's capacity are dropped: they skip the layer and pass through on the residual stream alone. Not a crash, just silent quality loss concentrated on the tokens the router was most confident about.
Underflow
Unused buffer slots are padded and computed anyway. Wasted FLOPs, wasted bandwidth.
Design choices that matter
Fine-grained experts. Rather than 8 experts the width of a dense FFN, use 256 narrow ones and route to 8. The active parameter count is unchanged, but the number of possible expert combinations explodes, so the model can express far more specialized mixtures at the same cost.
Shared experts. Reserve one or two experts that every token visits. Common knowledge that all tokens need — basic syntax, frequent patterns — lives there instead of being duplicated across every routed expert, which frees the routed ones to genuinely specialize.
Where to place MoE layers. Replacing every FFN is common, but many designs keep the first block or two dense, since early representations are still generic and routing on them is unreliable.
Router precision. Compute the router in FP32. It is a tiny matmul, and near-ties between experts in BF16 flip the top-$k$ selection, making routing non-deterministic across otherwise identical runs.
The honest cost sheet
| What you gain | What you pay |
|---|---|
| Many more parameters at fixed FLOPs per token | Memory scales with total parameters, so the model must be sharded across many devices regardless of how few are active |
| Better loss at equal training compute | An all-to-all dispatch and combine per MoE layer, the subject of the next note |
| Specialization across domains | A new failure mode — routing collapse — plus a hyperparameter ($\alpha$ or a bias schedule) that must be tuned |
| Cheaper inference per token than a dense model of equal capacity | Poor arithmetic intensity: each expert sees few tokens, so its matmuls are small and memory-bound |
Separate total from active parameters; route each token to $k$ of $E$ experts; keep the routing balanced with an auxiliary loss or a bias controller so the capacity factor can sit near one. Fine-grained experts widen the combination space and shared experts absorb what everyone needs.
Check yourself
- For $E=64$, $k=8$, give the total-to-active parameter ratio and the training FLOPs per token. calculation
- Explain why top-$k$ selection is non-differentiable and where the gradient does flow. reasoning
- Show that the auxiliary loss equals 1 under uniform routing for any $E$. derivation
- Compute the capacity for $T=8192$, $E=64$, $k=8$ at a capacity factor of 1.25. calculation
- Argue why fine-grained experts help even though the active parameter count is unchanged. analysis
- Explain why a bias-based balancer avoids perturbing the language-modelling gradient. design