The Transformer Decoder: From a Prefix to the Next Token
Start with the probability model
This chain-rule factorization does not itself require a Transformer. A causal Transformer is one way to parameterize each conditional distribution. With a BOS token, the representation at the BOS position predicts $x_1$; the representation at $x_1$ predicts $x_2$.
For input [BOS, red, kite], the last output distribution predicts what follows “kite.” It does not predict “kite” again. Keep this alignment explicit when slicing tensors or appending a generated token.
Follow the residual stream
Token and position information initialize $H^{(0)}\in\mathbb R^{n\times d}$. Each causal block mixes allowed prefix positions and applies a token-wise MLP, preserving residual width. After the final normalization, a vocabulary projection maps each row to $V$ logits. Softmax over that last axis defines token probabilities.
A common modern pre-norm decoder uses $U=H+\operatorname{CausalAttn}(\operatorname{Norm}(H))$ and $H'=U+\operatorname{MLP}(\operatorname{Norm}(U))$. Position handling, normalization type, activation, head sharing, and residual details vary. Learn the common graph, then inspect the exact checkpoint configuration.
Why causal training is parallel
During teacher-forced training, the entire observed sequence is already available. We can calculate all Q/K/V projections together and use a triangular mask so each prediction sees only its legal prefix. The targets do not have to be sampled by the current model first.
During ordinary generation, token $x_{t+1}$ is unknown until its distribution is computed and a decoding rule chooses it. The next forward step then consumes that chosen token. Parallel matrix computation within a step and sequential dependence between generated tokens are compatible facts.
Logits are not a decoding policy
Greedy decoding takes the highest-probability token. Temperature sampling uses $p_i\propto\exp(z_i/T)$ with $T>0$. Top-$k$ restricts support to a selected number of tokens; top-$p$ keeps a probability-mass prefix under a specified sorting and boundary convention. Renormalize the retained probabilities before sampling.
For logits $(2,1,0)$ at temperature 1, probabilities are approximately $(0.6652,0.2447,0.0900)$. Greedy always returns token 1; sampling sometimes returns another token. Lower temperature sharpens this particular distribution, but it is not a guarantee of factual accuracy.
The vocabulary softmax here is different from attention softmax. One selects a token distribution at the model output; the other mixes value vectors inside a layer. They normalize over different axes and play different roles.
What changes after the prompt?
Prefill computes representations for the prompt and saves each layer's causal keys and values. The final prompt position supplies the first continuation distribution. After choosing a token, process that token at its absolute position, append its new K/V entries at each layer, and read the combined prefix.
Processing one new query does not mean it attends to one key. It attends to the entire allowed cached prefix plus itself. If a prefix already contains $P$ tokens and a chunk contains $m$ new tokens, new query $i$ may read keys through absolute position $P+i$. An unshifted $m\times(P+m)$ triangular mask would be wrong.
The cache chapter derives the reuse argument and tests cached versus full-prefix outputs. Speculative decoding changes the execution schedule with a verification/correction algorithm.
The original decoder also has cross-attention
In the original encoder-decoder Transformer, a decoder layer includes causal target self-attention, source cross-attention, and an MLP. In a decoder-only language model, that separate encoder and cross-attention path are usually absent. Both are called decoders because their target-side computation supports autoregressive prediction, but their graphs differ.
See the original Transformer for the encoder-decoder graph and the GPT-2 report for a decoder-only language model.
Try it: When predicting the token after a prompt, do we need to feed an extra placeholder token first?
No. The logits at the final prompt position already predict the next token under the usual shifted causal objective. Feed the chosen continuation token when computing the following distribution.
A complete starting point
Continue to GPT-2 for a concrete parameter count and implementation path, then explore the costs of prefill and decoding.