Feed-Forward Networks: Expansion, Nonlinearity, and Gates
Attention moves information between positions; the feed-forward network transforms the features at each position. The feed-forward network is also called an MLP. It returns an update with the same width as its input, so the block can add that update to the running token vector.
Inside a block → the MLP branch. View the model diagram →
Transform one token vector at a time
Attention has already let “sat” read its context. The MLP now works on that position’s features. It uses the same learned weights at every position, but processes each vector separately. It does not fetch another token’s vector during this step.
A conventional MLP has three operations: expand the vector into more features, apply an activation to each feature, then project back to the original width. Here is a small example with two input features and three intermediate features.
(2, −1)One token vector(2, −1, 1)Make three combinations(2, 0, 1)ReLU sets negatives to zero(2, 1)Return two featuresFor the expansion above, choose three features: $x_1$, $x_2$, and $x_1+x_2$. For $x=(2,-1)$, those are $(2,-1,1)$. ReLU removes the negative entry. To project back, keep the first feature and sum the other two: $(2,0+1)=(2,1)$. These weights are chosen to make the arithmetic easy to follow.
Here is the same operation in matrix notation. The input width is $d$, the intermediate width is $m$, and $\sigma$ is the activation:
$W_1$ has shape $d\times m$ and expands the vector; $W_2$ has shape $m\times d$ and projects it back. The biases $b_1$ and $b_2$ shift the intermediate and output features. They are zero in our example.
Why the activation matters
Remove ReLU from the example and the middle vector stays $(2,-1,1)$. The same output projection now gives $(2,-1+1)=(2,0)$. The activation changed the computation by responding to the sign of the input features.
Without an activation, the two matrices could always collapse into a single matrix $W_1W_2$, plus a combined bias. Making the middle vector wider would still leave one linear transformation. The activation lets the MLP represent functions that a single linear map cannot.
“Position-wise” describes where the MLP reads from, not how much context its input contains. Its input may already encode information from many positions through earlier attention.
ReLU, GELU, and SiLU
ReLU discards negative inputs. GELU uses a Gaussian-CDF weighting; SiLU uses a sigmoid weighting. The latter two retain a smooth, small negative region rather than imposing a hard cutoff. Implementations may use an approximation for GELU, so exact numerical comparisons should record the chosen formula.
It is tempting to call each hidden coordinate a “fact” stored in a memory slot. That can be a useful hypothesis in a particular trained model, but the algebra only guarantees a learned nonlinear feature map. Features can be distributed, redundant, and dependent on context.
SwiGLU uses one feature to control another
A bias-free SwiGLU-style MLP computes
Both projected vectors have width $m$. The gate branch and the value branch are multiplied elementwise before the down-projection. This allows the strength of one learned feature to depend on another learned feature. The gate is not a normalized distribution over experts and need not lie between zero and one: SiLU can be negative and can exceed one.
The distinction matters when reading diagrams that label every multiplicative branch a “gate.” A sigmoid retention gate in a recurrent memory has a different range and job from a SiLU gate in a position-wise MLP. GLU Variants Improve Transformer develops these gated feed-forward alternatives.
Match parameter budgets before comparing widths
A conventional bias-free MLP has $2dm$ parameters. A gated MLP has three matrices and therefore $3dm$. Keeping $m=4d$ in both increases the gated MLP's matrix count from $8d^2$ to $12d^2$.
To match the conventional $m=4d$ budget, choose gated width $m\approx8d/3$, since $3d(8d/3)=8d^2$. For $d=768$, that gives $m=2048$ and 4,718,592 parameters in either design. Real implementations may round the width to a hardware-friendly multiple. Biases, expert routing, and shared matrices require separate accounting.
Keep the three projections visible in code
# x: [batch, sequence, d]; weights use row-vector convention.
def swiglu(x, w_gate, w_up, w_down):
gate_logits = x @ w_gate # [..., m]
gate = gate_logits / (1 + np.exp(-gate_logits))
value = x @ w_up # [..., m]
return (gate * value) @ w_down # [..., d]
This NumPy sketch exposes the operation; a robust implementation uses a numerically stable SiLU primitive rather than computing an overflowing exponential on extreme values. Fused kernels can combine operations and reduce intermediate memory traffic without changing the intended map.
Which cost grows with sequence length?
The MLP's matrix arithmetic is linear in token count at fixed widths: roughly $2ndm$ multiply-accumulate operations for the two-matrix version, or $3ndm$ for the gated version. Training may retain the expanded activations for backward, making intermediate width a memory decision too. Checkpointing can exchange recomputation for activation storage.
An MoE replaces one dense MLP with a routed collection of MLPs. Its experts often contain the same expansion-and-gating structure described here. The new questions are which experts execute and how their outputs and communication are handled; see Mixture of Experts.
Try it: A gated MLP with the same intermediate width as a two-matrix MLP has how many more matrix parameters?
Fifty percent more: $3dm$ instead of $2dm$. Comparing their accuracy without noting the changed parameter and compute budget can be misleading.
Prepare each branch’s input
Attention mixes context and the MLP transforms features. Next, normalization explains how the block prepares the vectors those branches receive. The complete block then puts both updates and residual additions together.