← All Posts

The Full GPT Model in C++ (LibTorch)

The Complete Picture

Token IDs (B, T) token_emb (B, T, n_embd) + Positional Encoding Transformer Block x N LN - MHA - +residual - LN - FFN - +residual (B, T, n_embd) in and out final LayerNorm lm_head - Logits (B, T, vocab_size)
  1. tokens — Raw integer IDs. Shape: (B, T) where B is batch size and T is sequence length.
  2. embedding — A lookup table converts each integer into a dense vector of dimension n_embd. Shape becomes (B, T, n_embd).
  3. positional encoding — Sinusoidal signals are added element-wise so the model can distinguish position 0 from position 5 from position 99. Shape stays (B, T, n_embd).
  4. [TransformerBlock] x N — Each block applies LayerNorm, multi-head attention with a residual connection, then LayerNorm, feed-forward with another residual connection. The shape stays (B, T, n_embd) through every block.
  5. final LayerNorm — One last normalization after all N blocks.
  6. lm_head — A linear projection from n_embd to vocab_size, producing logits. Final shape: (B, T, vocab_size).

The shape (B, T, n_embd) is the “communication bus” of the entire model. Every transformer block reads from it and writes back to it. Only the very first layer (embedding) and the very last layer (lm_head) change the last dimension.

The Seven Hyperparameters

The GPT constructor takes seven integers. Every architectural decision in the model follows from these seven numbers. Here is what each one controls:

ParameterMeaningTypical Value (GPT-2 Small)Affects
vocab_size Number of unique tokens the model can recognize. This is the size of the tokenizer's vocabulary. 50,257 token_emb rows, lm_head output dimension
n_embd Embedding dimension. The width of every hidden representation throughout the model. 768 Every layer's input/output width
num_heads Number of parallel attention heads in each transformer block. 12 Multi-head attention parallelism
head_size Dimension per attention head. Usually n_embd / num_heads. 64 Q, K, V projection sizes
d_ff Hidden dimension of the feed-forward network. Typically 4 * n_embd. 3,072 FFN expansion layer width
block_size Maximum sequence length the model can process. Also called context length or context window. 1,024 Positional encoding range, attention mask size
num_layers Number of transformer blocks stacked in sequence. The "depth" of the model. 12 Model depth, total parameter count
The relationship between head_size and n_embd. In most implementations, head_size = n_embd / num_heads. This ensures that concatenating all heads produces a vector of size num_heads * head_size = n_embd, which feeds cleanly into the next layer. Our code takes head_size as a separate parameter, giving you the flexibility to set it independently, but in practice you almost always want this relationship to hold.

Notice that block_size is stored but not directly used in the constructor or forward pass shown here. It is used elsewhere: the attention mask in each SingleHeadAttention module uses it to create the causal mask, and during inference it limits the context window. Storing it in the GPT class makes it accessible for generation loops.

Parameter scaling. Doubling num_layers roughly doubles the parameter count. Doubling n_embd roughly quadruples it (because every weight matrix has an n_embd dimension on both sides). This is why n_embd is the most powerful scaling knob.

The Constructor

Here is the complete constructor. It creates every submodule and registers each one with LibTorch's module system.

class GPT : public torch::nn::Module{ int vocab_size, n_embd, num_heads, head_size, d_ff, block_size, num_layers; shared_ptr<PositionalEncoding> pe; vector< shared_ptr<TransformerBlock> > transformer_blocks; torch::nn::Embedding token_emb{nullptr}; torch::nn::Linear lm_head{nullptr}; torch::nn::LayerNorm final_ln{nullptr}; public: GPT(int vocab_size, int n_embd, int num_heads, int head_size, int d_ff, int block_size, int num_layers) : vocab_size(vocab_size), n_embd(n_embd), num_heads(num_heads), head_size(head_size), d_ff(d_ff), block_size(block_size), num_layers(num_layers) { pe = make_shared<PositionalEncoding>(n_embd); lm_head = register_module( "lm_head", torch::nn::Linear( torch::nn::LinearOptions( n_embd, vocab_size ) ) ); final_ln = register_module("layer_norm", torch::nn::LayerNorm( torch::nn::LayerNormOptions( {n_embd} ) )); token_emb = register_module("token_emb", torch::nn::Embedding(vocab_size, n_embd) ); for (int layer_idx = 0; layer_idx < num_layers; layer_idx++){ shared_ptr<TransformerBlock> curr_block = make_shared<TransformerBlock>( num_heads, head_size, n_embd, d_ff ); register_module( "transformer_block_" + to_string(layer_idx), curr_block ); transformer_blocks.push_back(curr_block); } }

The Forward Pass

The forward method is remarkably compact. Six lines of code implement the entire GPT data flow:

torch::Tensor forward(torch::Tensor x){ x = token_emb(x); x = x + pe->forward(x); for (int layer_idx = 0; layer_idx < num_layers; layer_idx++){ x = transformer_blocks[layer_idx]->forward(x); } x = final_ln(x); x = lm_head(x); return x; }

Interactive Animation: The Forward Pass Pipeline

Click Step to advance through the forward pass, or Reset to start over. Watch how the shape and content transform at each stage.

Ready
Token IDs [ 5, 42, 13 ] shape: (1, 3) token_emb Lookup: each ID becomes a 64-dim vector shape: (1, 3, 64) + Positional Encoding Sinusoidal signals added to embeddings shape: (1, 3, 64) Transformer Block 0 LN, MHA, +residual, LN, FFN, +residual shape: (1, 3, 64) Transformer Block 1 LN, MHA, +residual, LN, FFN, +residual shape: (1, 3, 64) Final LayerNorm Normalize to mean=0, var=1, then scale+shift shape: (1, 3, 64) lm_head (Linear) Project n_embd to vocab_size: logits over vocabulary shape: (1, 3, vocab_size) Output Logits Ready for cross-entropy loss or sampling

Why Embedding Has No √d Scaling

If you have read the original "Attention Is All You Need" paper, you might recall this line:

In the embedding layers, we multiply those weights by sqrt(d_model).

The original Transformer multiplies the embedding output by √n_embd before adding positional encoding. The reasoning: embedding vectors tend to have small magnitudes (because they are initialized randomly with small values), while sinusoidal positional encodings have magnitudes around 1.0. Multiplying by √n_embd scales the embeddings up so that token identity is not drowned out by positional signals.

Our GPT class does not include this scaling:

x = token_emb(x); x = x + pe->forward(x); // no sqrt(n_embd) multiplier

This omission is intentional and common in GPT-style models. Here is why it works fine:

  • Learned embeddings compensate. During training, the embedding vectors naturally adjust their magnitudes. If the model needs larger embeddings to balance against positional encoding, gradient descent will make them larger.
  • LayerNorm normalizes anyway. The very first operation inside each transformer block is LayerNorm, which normalizes the input to zero mean and unit variance. Any magnitude differences between embeddings and positional encodings get washed out immediately.
  • GPT-2 and GPT-3 skip it. The original GPT paper (Radford et al.) does not use this scaling, and subsequent models followed suit. It is an encoder-side convention from the original Transformer that decoder-only models dropped.
If you add it: Including the √n_embd scaling will not break anything. It just changes the initial magnitude of embeddings relative to positional encodings. The model will compensate during training. You would write: x = token_emb(x) * sqrt(n_embd); then add PE.

Shape Trace Table

Here is the complete shape trace from input to output. We use concrete numbers: B=4 (batch size), T=32 (sequence length), n_embd=64, vocab_size=1000, num_layers=2.

#OperationInput ShapeOutput ShapeNotes
0 Function input x - (4, 32) Integer tensor, dtype=long
1 token_emb(x) (4, 32) (4, 32, 64) Lookup: each int becomes a 64-dim vector
2a pe->forward(x) (4, 32, 64) (32, 64) PE output; broadcasts across batch
2b x + pe->forward(x) (4,32,64) + (32,64) (4, 32, 64) Broadcast addition
3 transformer_blocks[0]->forward(x) (4, 32, 64) (4, 32, 64) Block 0: LN, MHA, +res, LN, FFN, +res
4 transformer_blocks[1]->forward(x) (4, 32, 64) (4, 32, 64) Block 1: same structure, different weights
5 final_ln(x) (4, 32, 64) (4, 32, 64) Normalize last dim, learned scale+shift
6 lm_head(x) (4, 32, 64) (4, 32, 1000) Linear: 64 to 1000 (vocab_size)

The shape (4, 32, 64) persists through six of eight operations. Only the embedding (adding the last dimension) and lm_head (changing the last dimension) break this pattern. This uniformity is a core design principle of the Transformer: every internal representation lives in the same vector space.

Why does shape preservation matter? It enables stacking. If block k produced a different shape than what block k+1 expects, you would need adapter layers between every pair of blocks. Shape preservation means you can stack 2 blocks or 200 blocks with zero architectural changes. The only thing that changes is the num_layers hyperparameter.

Parameter Count

Let us count every learnable parameter in the model. We will use two sets of numbers: symbolic formulas and concrete values for a "small" configuration.

Example configuration: vocab_size=1000, n_embd=64, num_heads=4, head_size=16, d_ff=256, block_size=128, num_layers=2.

Token Embedding

ComponentFormulaCount
Embedding weightvocab_size * n_embd1,000 * 64 = 64,000

One Transformer Block

Each block contains a multi-head attention module, a feed-forward module, and two LayerNorm layers.

Multi-Head Attention (MHA)

MHA consists of num_heads single attention heads. Each head has three linear projections (Q, K, V) of size (n_embd, head_size). There is no output projection in our implementation (the concatenated heads already have dimension num_heads * head_size = n_embd).

ComponentFormulaCount
Q weight per headn_embd * head_size64 * 16 = 1,024
Q bias per headhead_size16
K weight per headn_embd * head_size1,024
K bias per headhead_size16
V weight per headn_embd * head_size1,024
V bias per headhead_size16
Per head total3 * (n_embd * head_size + head_size)3,120
All heads totalnum_heads * 3 * (n_embd * head_size + head_size)4 * 3,120 = 12,480

Feed-Forward Network (FFN)

ComponentFormulaCount
Linear 1 weightn_embd * d_ff64 * 256 = 16,384
Linear 1 biasd_ff256
Linear 2 weightd_ff * n_embd256 * 64 = 16,384
Linear 2 biasn_embd64
FFN total2 * n_embd * d_ff + d_ff + n_embd33,088

Two LayerNorms per block

ComponentFormulaCount
LN1 scale (gamma)n_embd64
LN1 shift (beta)n_embd64
LN2 scale + shift2 * n_embd128
Both LNs total4 * n_embd256

One block total

ComponentCount
MHA12,480
FFN33,088
2 x LayerNorm256
Block total45,824

Final LayerNorm

ComponentFormulaCount
Scale (gamma)n_embd64
Shift (beta)n_embd64
Total2 * n_embd128

Language Model Head (lm_head)

ComponentFormulaCount
Weightn_embd * vocab_size64 * 1,000 = 64,000
Biasvocab_size1,000
Totaln_embd * vocab_size + vocab_size65,000

Grand Total

ComponentCount% of Total
token_emb64,00026.8%
Block 045,82419.2%
Block 145,82419.2%
final_ln1280.1%
lm_head65,00027.2%
Grand total220,776-

A few observations:

  • Embeddings dominate at small scale. The token_emb and lm_head together account for about 54% of all parameters. This is because vocab_size (1,000) is large relative to n_embd (64).
  • At GPT-2 scale, blocks dominate. With vocab_size=50,257 and n_embd=768 and 12 layers, the embedding and lm_head account for about 50M parameters total, while the 12 blocks account for roughly 85M. As models get deeper, the block parameters grow proportionally.
  • LayerNorm is tiny. Each LayerNorm adds only 2*n_embd parameters. Even with dozens of them, they contribute less than 1% of total parameters.
  • FFN is the largest part of each block. The feed-forward network (33,088) is larger than multi-head attention (12,480) inside each block. This is because d_ff = 4 * n_embd creates a wide expansion layer.
Weight tying. Some models (including GPT-2) tie the token_emb weight matrix to the lm_head weight matrix. Since both have shape (vocab_size, n_embd) (the linear layer stores weights transposed), sharing them eliminates one copy. This cuts roughly vocab_size * n_embd parameters. Our implementation does not use weight tying, so both are independent.

General Formula

Total = vocab_size * n_embd (token_emb) + num_layers * [ num_heads * 3 * (n_embd * head_size + head_size) (MHA) + 2 * n_embd * d_ff + d_ff + n_embd (FFN) + 4 * n_embd (2 LayerNorms) ] + 2 * n_embd (final_ln) + n_embd * vocab_size + vocab_size (lm_head)