Mixture of Experts: Conditional MLPs and the Cost of Routing
Start with the dense MLP
A dense feed-forward layer applies the same function $E(x)$ independently to every token representation $x\in\mathbb R^d$. A mixture of experts stores $E_1,\ldots,E_N$ and chooses a small subset for each token. These are neural functions, often gated MLPs, rather than separate complete language models.
A router computes scores $s=W_rx$ and selects a set $\mathcal T(x)=\operatorname{TopK}(s,k)$. One possible convention normalizes softmax weights over the selected set:
This is an explicitly chosen reference rule. Implementations differ: some use sigmoid scores, retain weights normalized over all experts, add shared experts, or apply router biases. The formula must match the implementation before comparing output scales or reproducing a checkpoint.
Route one two-dimensional token
Suppose four expert scores are $(2,1,0,-1)$ and $k=2$. The selected experts are 1 and 2. Their renormalized weights are approximately $(0.7311,0.2689)$. If their outputs are $E_1(x)=(2,0)$ and $E_2(x)=(0,4)$, the mixture is $(1.4621,1.0758)$.
The router chooses the functions using the current representation; the experts then transform that representation. Expert names such as “math expert” are not supplied by this equation. A specialization claim needs evidence about learned routing and behavior.
Total capacity and active computation
If each expert contains $P_E$ parameters, a routed bank stores $NP_E$ expert parameters but activates approximately $kP_E$ per token. Add the router, attention layers, embeddings, shared experts, and other dense components to count a complete model. If $N=64$ and $k=2$, the ratio of stored to selected expert parameters is 32; that does not make the whole model 32 times cheaper than a particular dense baseline.
For a three-matrix gated expert with input width $d$ and intermediate width $m$, $P_E\approx3dm$ without biases. Its main forward matrix multiplications use roughly $3dm$ multiply-accumulates per token. Routing and dispatch have additional costs. Matching active parameters does not automatically match wall time, training compute, or quality.
Follow a batch through expert parallelism
Flatten batch and sequence positions into token rows. Compute selected expert IDs and weights; group rows by destination expert; send them to the devices that own those experts; execute each expert on its received batch; return outputs; then scatter and combine them into original token order.
The communication pattern is commonly an all-to-all exchange when experts are distributed. A device can be idle while another processes an overloaded expert. Tiny expert batches may underuse matrix hardware even with little nominal arithmetic. The router's statistical decisions therefore create a systems workload.
When one token goes to multiple experts, track a separate dispatch entry and weight for each assignment. The combine operation sums contributions; overwriting the output with the last returning expert is a subtle implementation error.
Why load balancing appears in the training objective
A router can collapse onto a few experts. One family of remedies penalizes imbalance using observed expert frequencies and router probabilities. For a top-1 illustration, with $f_i$ the fraction of tokens assigned to expert $i$ and $P_i$ the batch mean router probability, an auxiliary term proportional to $N\sum_i f_iP_i$ encourages balanced use. The hard assignment counts themselves are not smoothly differentiable.
Capacity limits cap the number of token assignments accepted by each expert. Overflow can be dropped, rerouted, or handled with a different dispatch policy. Those choices affect the actual function and require explicit reporting. Dropless systems avoid token dropping but still have to manage imbalanced work and memory.
The foundational references are the sparsely gated MoE layer and Switch Transformers. Their routing and balancing choices should not be treated as universal MoE requirements.
Shared and routed experts
A shared expert runs on every token and contributes alongside the selected experts. It can provide a common transformation while routed experts supply conditional capacity. Algebraically, $y=E_{\mathrm{shared}}(x)+\sum_{i\in\mathcal T}g_iE_i(x)$, possibly with additional gates or multiple shared experts. See DeepSeekMoE for one concrete shared/routed design.
Move routed computation into a narrower space
A further design choice compresses the routed path: $z=W_\downarrow x\in\mathbb R^{d_\ell}$, computes the routed mixture in that space, and maps it back with $W_\uparrow$. Shared experts can remain in the original width:
This illustrative decomposition separates dimensional compression from sparse selection. The down/up projections cost computation, and the latent bottleneck restricts the routed path. The shared path and residual stream can still carry full-width information. Kimi K3 provides a specific latent-MoE configuration.
Try it: Does top-2 routing mean that exactly two experts run for an entire batch?
No. It means two routed experts per token. Different tokens may choose different pairs, so a batch can exercise the whole bank. Shared experts, if present, run in addition.
Keep the design axes separate
Read Attention Residuals for routing across depth, and the comparison map for how MoE combines with attention, recurrence, and cache compression.