The Full GPT Model in C++ (LibTorch)
The Complete Picture
- tokens — Raw integer IDs. Shape:
(B, T)where B is batch size and T is sequence length. - embedding — A lookup table converts each integer into a dense vector of dimension
n_embd. Shape becomes(B, T, n_embd). - positional encoding — Sinusoidal signals are added element-wise so the model can distinguish position 0 from position 5 from position 99. Shape stays
(B, T, n_embd). - [TransformerBlock] x N — Each block applies LayerNorm, multi-head attention with a residual connection, then LayerNorm, feed-forward with another residual connection. The shape stays
(B, T, n_embd)through every block. - final LayerNorm — One last normalization after all N blocks.
- lm_head — A linear projection from
n_embdtovocab_size, producing logits. Final shape:(B, T, vocab_size).
The shape (B, T, n_embd) is the “communication bus” of the entire model. Every transformer block reads from it and writes back to it. Only the very first layer (embedding) and the very last layer (lm_head) change the last dimension.
The Seven Hyperparameters
The GPT constructor takes seven integers. Every architectural decision in the model follows from these seven numbers. Here is what each one controls:
| Parameter | Meaning | Typical Value (GPT-2 Small) | Affects |
|---|---|---|---|
vocab_size |
Number of unique tokens the model can recognize. This is the size of the tokenizer's vocabulary. | 50,257 | token_emb rows, lm_head output dimension |
n_embd |
Embedding dimension. The width of every hidden representation throughout the model. | 768 | Every layer's input/output width |
num_heads |
Number of parallel attention heads in each transformer block. | 12 | Multi-head attention parallelism |
head_size |
Dimension per attention head. Usually n_embd / num_heads. |
64 | Q, K, V projection sizes |
d_ff |
Hidden dimension of the feed-forward network. Typically 4 * n_embd. |
3,072 | FFN expansion layer width |
block_size |
Maximum sequence length the model can process. Also called context length or context window. | 1,024 | Positional encoding range, attention mask size |
num_layers |
Number of transformer blocks stacked in sequence. The "depth" of the model. | 12 | Model depth, total parameter count |
head_size = n_embd / num_heads. This ensures that concatenating all heads produces a vector of size num_heads * head_size = n_embd, which feeds cleanly into the next layer. Our code takes head_size as a separate parameter, giving you the flexibility to set it independently, but in practice you almost always want this relationship to hold.
Notice that block_size is stored but not directly used in the constructor or forward pass shown here. It is used elsewhere: the attention mask in each SingleHeadAttention module uses it to create the causal mask, and during inference it limits the context window. Storing it in the GPT class makes it accessible for generation loops.
num_layers roughly doubles the parameter count. Doubling n_embd roughly quadruples it (because every weight matrix has an n_embd dimension on both sides). This is why n_embd is the most powerful scaling knob.
The Constructor
Here is the complete constructor. It creates every submodule and registers each one with LibTorch's module system.
The Forward Pass
The forward method is remarkably compact. Six lines of code implement the entire GPT data flow:
Interactive Animation: The Forward Pass Pipeline
Click Step to advance through the forward pass, or Reset to start over. Watch how the shape and content transform at each stage.
Why Embedding Has No √d Scaling
If you have read the original "Attention Is All You Need" paper, you might recall this line:
The original Transformer multiplies the embedding output by √n_embd before adding positional encoding. The reasoning: embedding vectors tend to have small magnitudes (because they are initialized randomly with small values), while sinusoidal positional encodings have magnitudes around 1.0. Multiplying by √n_embd scales the embeddings up so that token identity is not drowned out by positional signals.
Our GPT class does not include this scaling:
This omission is intentional and common in GPT-style models. Here is why it works fine:
- Learned embeddings compensate. During training, the embedding vectors naturally adjust their magnitudes. If the model needs larger embeddings to balance against positional encoding, gradient descent will make them larger.
- LayerNorm normalizes anyway. The very first operation inside each transformer block is LayerNorm, which normalizes the input to zero mean and unit variance. Any magnitude differences between embeddings and positional encodings get washed out immediately.
- GPT-2 and GPT-3 skip it. The original GPT paper (Radford et al.) does not use this scaling, and subsequent models followed suit. It is an encoder-side convention from the original Transformer that decoder-only models dropped.
√n_embd scaling will not break anything. It just changes the initial magnitude of embeddings relative to positional encodings. The model will compensate during training. You would write: x = token_emb(x) * sqrt(n_embd); then add PE.
Shape Trace Table
Here is the complete shape trace from input to output. We use concrete numbers: B=4 (batch size), T=32 (sequence length), n_embd=64, vocab_size=1000, num_layers=2.
| # | Operation | Input Shape | Output Shape | Notes |
|---|---|---|---|---|
| 0 | Function input x |
- | (4, 32) |
Integer tensor, dtype=long |
| 1 | token_emb(x) |
(4, 32) |
(4, 32, 64) |
Lookup: each int becomes a 64-dim vector |
| 2a | pe->forward(x) |
(4, 32, 64) |
(32, 64) |
PE output; broadcasts across batch |
| 2b | x + pe->forward(x) |
(4,32,64) + (32,64) |
(4, 32, 64) |
Broadcast addition |
| 3 | transformer_blocks[0]->forward(x) |
(4, 32, 64) |
(4, 32, 64) |
Block 0: LN, MHA, +res, LN, FFN, +res |
| 4 | transformer_blocks[1]->forward(x) |
(4, 32, 64) |
(4, 32, 64) |
Block 1: same structure, different weights |
| 5 | final_ln(x) |
(4, 32, 64) |
(4, 32, 64) |
Normalize last dim, learned scale+shift |
| 6 | lm_head(x) |
(4, 32, 64) |
(4, 32, 1000) |
Linear: 64 to 1000 (vocab_size) |
The shape (4, 32, 64) persists through six of eight operations. Only the embedding (adding the last dimension) and lm_head (changing the last dimension) break this pattern. This uniformity is a core design principle of the Transformer: every internal representation lives in the same vector space.
Parameter Count
Let us count every learnable parameter in the model. We will use two sets of numbers: symbolic formulas and concrete values for a "small" configuration.
vocab_size=1000, n_embd=64, num_heads=4, head_size=16, d_ff=256, block_size=128, num_layers=2.
Token Embedding
| Component | Formula | Count |
|---|---|---|
| Embedding weight | vocab_size * n_embd | 1,000 * 64 = 64,000 |
One Transformer Block
Each block contains a multi-head attention module, a feed-forward module, and two LayerNorm layers.
Multi-Head Attention (MHA)
MHA consists of num_heads single attention heads. Each head has three linear projections (Q, K, V) of size (n_embd, head_size). There is no output projection in our implementation (the concatenated heads already have dimension num_heads * head_size = n_embd).
| Component | Formula | Count |
|---|---|---|
| Q weight per head | n_embd * head_size | 64 * 16 = 1,024 |
| Q bias per head | head_size | 16 |
| K weight per head | n_embd * head_size | 1,024 |
| K bias per head | head_size | 16 |
| V weight per head | n_embd * head_size | 1,024 |
| V bias per head | head_size | 16 |
| Per head total | 3 * (n_embd * head_size + head_size) | 3,120 |
| All heads total | num_heads * 3 * (n_embd * head_size + head_size) | 4 * 3,120 = 12,480 |
Feed-Forward Network (FFN)
| Component | Formula | Count |
|---|---|---|
| Linear 1 weight | n_embd * d_ff | 64 * 256 = 16,384 |
| Linear 1 bias | d_ff | 256 |
| Linear 2 weight | d_ff * n_embd | 256 * 64 = 16,384 |
| Linear 2 bias | n_embd | 64 |
| FFN total | 2 * n_embd * d_ff + d_ff + n_embd | 33,088 |
Two LayerNorms per block
| Component | Formula | Count |
|---|---|---|
| LN1 scale (gamma) | n_embd | 64 |
| LN1 shift (beta) | n_embd | 64 |
| LN2 scale + shift | 2 * n_embd | 128 |
| Both LNs total | 4 * n_embd | 256 |
One block total
| Component | Count |
|---|---|
| MHA | 12,480 |
| FFN | 33,088 |
| 2 x LayerNorm | 256 |
| Block total | 45,824 |
Final LayerNorm
| Component | Formula | Count |
|---|---|---|
| Scale (gamma) | n_embd | 64 |
| Shift (beta) | n_embd | 64 |
| Total | 2 * n_embd | 128 |
Language Model Head (lm_head)
| Component | Formula | Count |
|---|---|---|
| Weight | n_embd * vocab_size | 64 * 1,000 = 64,000 |
| Bias | vocab_size | 1,000 |
| Total | n_embd * vocab_size + vocab_size | 65,000 |
Grand Total
| Component | Count | % of Total |
|---|---|---|
| token_emb | 64,000 | 26.8% |
| Block 0 | 45,824 | 19.2% |
| Block 1 | 45,824 | 19.2% |
| final_ln | 128 | 0.1% |
| lm_head | 65,000 | 27.2% |
| Grand total | 220,776 | - |
A few observations:
- Embeddings dominate at small scale. The token_emb and lm_head together account for about 54% of all parameters. This is because vocab_size (1,000) is large relative to n_embd (64).
- At GPT-2 scale, blocks dominate. With vocab_size=50,257 and n_embd=768 and 12 layers, the embedding and lm_head account for about 50M parameters total, while the 12 blocks account for roughly 85M. As models get deeper, the block parameters grow proportionally.
- LayerNorm is tiny. Each LayerNorm adds only 2*n_embd parameters. Even with dozens of them, they contribute less than 1% of total parameters.
- FFN is the largest part of each block. The feed-forward network (33,088) is larger than multi-head attention (12,480) inside each block. This is because d_ff = 4 * n_embd creates a wide expansion layer.
(vocab_size, n_embd) (the linear layer stores weights transposed), sharing them eliminates one copy. This cuts roughly vocab_size * n_embd parameters. Our implementation does not use weight tying, so both are independent.