← All Posts
Deep Learning · Transformers· Architecture and model families

The Transformer Encoder: Context, Padding, and Task Heads

An encoder turns an observed sequence into contextual vectors. Its job is not inherently to generate the next token. The output head and training objective decide how those vectors are used.

Track one batch through the network

For a batch of $B$ sequences padded to length $n$, token IDs have shape $B\times n$. Embedding lookup produces $B\times n\times d$. Add or otherwise incorporate the model's position representation, then pass the tensor through a stack of self-attention and feed-forward blocks.

In a pre-normalized encoder, one block can be written $U=X+\operatorname{Attn}(\operatorname{Norm}(X))$ and $Y=U+\operatorname{MLP}(\operatorname{Norm}(U))$. A post-normalized encoder instead normalizes after each residual addition. These are different graphs; the original Transformer and BERT should not silently be rewritten as modern pre-norm models when reproducing their checkpoints.

Bidirectional means available context, not two separate passes

For an unpadded sequence of four positions, every query can read every key. There is no need to run one left-to-right and another right-to-left recurrence. The attention matrix directly represents all permitted pairs. The representation at position 2 can therefore change when position 4 changes.

This is useful for disambiguating an observed word from its complete sentence. It also means the ordinary causal KV-cache proof fails if new text is appended: earlier contextual representations can change, so cached upper-layer keys and values are no longer guaranteed valid.

A padding key must not influence a valid token

Suppose sequence A has length 4 and sequence B has length 2, both stored in four slots. In B, valid queries should exclude key slots 3 and 4. Merely assigning a zero embedding to padding is insufficient: projections, position terms, and softmax normalization can still make those slots affect the output.

Masking padding keys does not necessarily force padded query outputs to zero. You may ignore them in pooling and loss, or explicitly zero them according to the model's convention. Keep at least one allowed key for every query whose output is actually computed with ordinary softmax; fully masked rows need defined handling.

Three heads for three tasks

For token labeling, apply a shared linear map to each contextual vector: $Z=HW_{\mathrm{tag}}+b$, with shape $B\times n\times C$ for $C$ labels. Ignore padding and any unlabelled subtokens according to the annotation policy.

For sequence classification, construct one sequence vector. A dedicated classification token is one option; a masked mean is another:

$$h_{\mathrm{pool}}=\frac{\sum_i m_i h_i}{\sum_i m_i}.$$

If two valid outputs are $(2,0)$ and $(0,4)$, the masked mean is $(1,2)$. Averaging over four stored slots, including two zeros, gives $(0.5,1)$ instead. A classifier can be sensitive to that unintended dependence on padding length.

For masked-token prediction, a vocabulary head produces logits at selected positions. The input must hide or alter the information the task asks the model to reconstruct. A full-attention encoder given the unchanged target token has an easy shortcut.

A useful embedding is a trained behavior

Taking the mean of any encoder's hidden states does not automatically create a high-quality semantic retrieval embedding. The pooling method, objective, normalization, and similarity function matter. A classification encoder and a contrastively trained retrieval encoder can share the same backbone but organize sequence vectors differently.

Likewise, an attention heatmap is an internal weighting visualization, not a complete explanation of a classifier. Residual paths, value projections, feed-forward functions, and later layers all contribute.

Verify the information graph

Change a valid later token and check that an earlier output can respond; this distinguishes bidirectional behavior from an accidental causal mask. Add extra padding and confirm valid outputs and masked pooling remain consistent. Compare batched and unbatched execution in evaluation mode, accounting for floating-point tolerance.

When fine-tuning, distinguish freezing parameters from switching to evaluation mode. Frozen parameters do not receive updates, but dropout can still be active if the model remains in training mode. Choose both settings intentionally.

Try it: If padded embeddings are all zeros, may we omit the padding mask?

No. Even zero-valued keys can receive softmax probability and change normalization, and position terms or projection biases may make them nonzero. Mask invalid keys explicitly.

The encoder as a pretraining model

Read BERT for the corruption objective and original input format. For the original encoder-decoder construction, see Attention Is All You Need and cross-attention.