Dropout: What Is Random, What Is Preserved
Dropout randomly zeros some activations during training. It scales up the survivors so their average value is preserved. At evaluation, ordinary inverted dropout passes the input through unchanged.
On the update branches → optional training noise. View the model diagram →
Follow one number through dropout
Suppose one feature has value $3$ and the drop probability is $p=0.25$. On each training pass, dropout makes one of two choices:
| Outcome | Probability | Output |
|---|---|---|
| Drop the feature | 25% | $0$ |
| Keep and scale it | 75% | $3/(1-0.25)=4$ |
Across repeated draws, the mean is $0.25\times0+0.75\times4=3$. The output of an individual draw is still zero or four. Scaling the survivors preserves the average at this point in the computation, not every individual activation.
This is inverted dropout. In general, use a keep flag $r$ that equals 1 with probability $1-p$ and 0 otherwise:
If the drop probability is zero, every value passes through unchanged. The formula is undefined at $p=1$; an implementation that supports that endpoint needs a separate rule.
Optional: how much noise does dropout add?
The keep flag is a Bernoulli random variable: $r\sim\operatorname{Bernoulli}(1-p)$. Conditional on the current activation $x$,
Our example has variance $3$. Increasing $p$ makes the surviving values larger and the draws noisier, while preserving the same mean.
Why add this noise?
Training with randomly missing activations can discourage brittle reliance on a particular combination of features. The optimization objective effectively averages over sampled masks. This is a regularization mechanism, not a guarantee that every feature becomes independent or that the network cannot overfit. The original reference is Dropout: A Simple Way to Prevent Neural Networks from Overfitting.
The expectation calculation is local. In general, $\mathbb E[f(y)]\ne f(\mathbb E[y])$ for a nonlinear $f$. Therefore the deterministic evaluation network is not exactly the arithmetic average of every possible masked network's predictions. The common “ensemble” intuition is approximate.
Where does dropout sit in a transformer?
Embedding dropout perturbs input representations. MLP dropout perturbs intermediate features or the branch output. Residual-branch dropout perturbs an update before adding it to the residual stream. These positions affect both the distribution of activations and the gradients.
Mask shape also matters. Independent masks for each token and feature differ from a mask shared across positions. The latter removes a feature consistently across a sequence. Neither should be silently substituted for the other when reproducing a model.
Attention dropout comes after softmax
A causal mask removes forbidden positions before softmax. Attention dropout instead perturbs the weights after they have been normalized. It serves a different purpose and does not make the attention causal.
Suppose attention probabilities are $a=(0.2,0.3,0.5)$ and the last entry is dropped under $p=0.5$. If the first two survive, inverted dropout gives $(0.4,0.6,0)$, which happens to sum to one. If only the last survives, the sum is also one; but if only the first survives, the row becomes $(0.4,0,0)$ and sums to $0.4$. Other masks give other sums.
Conventional attention dropout applies the Bernoulli mask after softmax and rescales surviving entries without renormalizing the row. Each probability is preserved in expectation. A single sampled output need not be a convex combination of values. Dropping logits before softmax or renormalizing survivors implements a different operation.
Evaluation mode is part of correctness
With inverted dropout, evaluation usually returns the input unchanged. If dropout remains active, repeated inference calls can differ and cached/full-sequence comparisons may fail for reasons unrelated to caching. Conversely, accidentally disabling dropout during training changes the regularized objective.
def dropout(x, p, training, rng):
if not training or p == 0:
return x
if not 0 <= p < 1:
raise ValueError("drop probability must be in [0, 1)")
keep = rng.random(x.shape) >= p
return x * keep / (1 - p)
The random generator is an explicit argument so a test can control it. For reproducibility, the seed alone may not be enough: changing batch shape, distributed partitioning, or kernel implementation can change which random draws map to which activations.
Interpret a dropout setting in its training context
There is no universal best drop rate for transformers. Dataset size, training duration, model capacity, other regularization, and the target task all affect the tradeoff. A checkpoint trained with zero dropout is not necessarily missing a component; adding dropout at inference does not recreate a training regularizer it never learned under.
Try it: Does setting dropout to zero turn off the residual connection?
No. Dropout is the identity when its drop probability is zero. The residual addition still sums the branch update with the stream. Turning a branch off would require setting its output to zero.
Give the model a sense of position
You have now seen the main operations in one transformer block. Continue with position information to understand how the model can distinguish where tokens occur. The training chapter puts randomness, loss, and evaluation mode together.