← All Posts
Deep Learning · Popular Videos · Umar Jamil· Part 13 · Chapter 26

Mixture of Experts from First Principles

Capacity and cost are the same thing in a dense model, and they need not be. Every token pays for every parameter because every parameter participates in every forward pass. Mixture of experts breaks that link: hold many parameters, activate a few per token, and pay only for what you activate.

Breaking the link

From the FLOPs note, a dense model costs $6N$ training FLOPs per token, where $N$ counts all parameters. Scaling capacity therefore scales cost one-for-one.

An MoE layer replaces the single feed-forward network with $E$ independent copies — the experts — plus a small router that picks which $k$ of them each token visits.

$$\text{FFN}(\mathbf x)\ \longrightarrow\ \sum_{i\in\text{Top-}k(\mathbf x)}g_i(\mathbf x)\,E_i(\mathbf x).$$
Two parameter counts, not one. Total parameters set memory; active parameters set FLOPs per token. With $E$ experts and top-$k$ routing the ratio is $E/k$, and only the active count enters $6N$.
ConfigurationExperts $E$Top-$k$Total FFN paramsActive per tokenSparsity
Dense11baselinesame100%
8 large experts8245.1B11.3B25%
256 fine experts + 1 shared2568656.5B23.0B3.5%

The second row buys 29 times the parameters for the FLOPs of a much smaller model. That is the entire pitch, and the rest of this note is about what it costs.

The router

The router is a single linear layer, $E$ outputs wide:

$$\mathbf s=\text{softmax}\big(W_r\mathbf x\big)\in\mathbb R^{E},\qquad \text{Top-}k(\mathbf x)=\text{indices of the }k\text{ largest }s_i,$$

and the chosen experts' outputs are combined with their (renormalized) gate values as weights. It is a tiny amount of arithmetic guarding an enormous amount of parameter.

Top-$k$ is not differentiable, and that is the central difficulty. The gradient flows through the gate values $g_i$ of the selected experts, never through the selection itself. An expert that is never chosen receives no gradient, so it never improves, so it is never chosen. Routing collapse is the default outcome unless something actively prevents it.

Load balancing

The standard remedy is an auxiliary loss that penalizes uneven routing. Let $f_i$ be the fraction of tokens dispatched to expert $i$ and $P_i$ the mean router probability assigned to it:

$$\mathcal L_{\text{aux}}=\alpha\,E\sum_{i=1}^{E}f_i\,P_i.$$

The $f_i$ term is a count and carries no gradient; $P_i$ does. Pushing down $P_i$ for over-subscribed experts is what nudges the router toward balance. The multiplier $E$ normalizes the loss so that a perfectly uniform distribution gives $\mathcal L_{\text{aux}}=1$ for any $E$, which makes $\alpha$ transferable across configurations.

$\alpha$ too small

Experts collapse onto a few favourites. Effective capacity drops to a fraction of what you paid for, and the extra parameters are dead weight.

$\alpha$ too large

Routing becomes near-uniform, which means near-random. Experts cannot specialize, and the model behaves like a noisier dense one.

Because the auxiliary loss perturbs the objective you actually care about, recent systems prefer a bias-based alternative: keep a per-expert bias added to the routing scores only for the top-$k$ selection, and adjust it up or down between steps according to observed load. Balance is enforced without touching the gradient of the language-modelling loss at all.

Capacity and dropped tokens

Balance matters for a second, purely mechanical reason. Expert buffers must be allocated with a fixed shape before routing is known, so each expert gets room for a bounded number of tokens:

$$C=\text{capacity factor}\times\frac{k\,T}{E},$$

where $T$ is the tokens in the batch and $kT/E$ is the average load. A capacity factor of $1.25$ gives 25% headroom.

Overflow

Tokens beyond an expert's capacity are dropped: they skip the layer and pass through on the residual stream alone. Not a crash, just silent quality loss concentrated on the tokens the router was most confident about.

Underflow

Unused buffer slots are padded and computed anyway. Wasted FLOPs, wasted bandwidth.

The capacity factor is a direct trade between dropped tokens and wasted compute, and good load balancing is what lets you set it near 1.0. It also matters enormously for the all-to-all in the next note, which needs fixed-size buffers to be efficient.
total FFN params
active per token
tokens dropped
buffer wasted
A $d=7168$ model with expert width 2048 over 58 layers, batch of 8192 tokens. “Routing skew” interpolates from perfectly uniform routing to a heavily collapsed distribution.

Design choices that matter

1

Fine-grained experts. Rather than 8 experts the width of a dense FFN, use 256 narrow ones and route to 8. The active parameter count is unchanged, but the number of possible expert combinations explodes, so the model can express far more specialized mixtures at the same cost.

2

Shared experts. Reserve one or two experts that every token visits. Common knowledge that all tokens need — basic syntax, frequent patterns — lives there instead of being duplicated across every routed expert, which frees the routed ones to genuinely specialize.

3

Where to place MoE layers. Replacing every FFN is common, but many designs keep the first block or two dense, since early representations are still generic and routing on them is unreliable.

4

Router precision. Compute the router in FP32. It is a tiny matmul, and near-ties between experts in BF16 flip the top-$k$ selection, making routing non-deterministic across otherwise identical runs.

The honest cost sheet

What you gainWhat you pay
Many more parameters at fixed FLOPs per tokenMemory scales with total parameters, so the model must be sharded across many devices regardless of how few are active
Better loss at equal training computeAn all-to-all dispatch and combine per MoE layer, the subject of the next note
Specialization across domainsA new failure mode — routing collapse — plus a hyperparameter ($\alpha$ or a bias schedule) that must be tuned
Cheaper inference per token than a dense model of equal capacityPoor arithmetic intensity: each expert sees few tokens, so its matmuls are small and memory-bound
MoE is a systems bet as much as a modelling one. The parameters have to live somewhere, and unlike a dense model you cannot simply replicate them. That is why expert parallelism exists, and why MoE only pays off at a scale where you were going to shard anyway.
Takeaway

Separate total from active parameters; route each token to $k$ of $E$ experts; keep the routing balanced with an auxiliary loss or a bias controller so the capacity factor can sit near one. Fine-grained experts widen the combination space and shared experts absorb what everyone needs.

Check yourself