TransformerBlock in C++ (LibTorch)
We have built multi-head attention and the feed-forward network in previous posts. Now we combine them into the TransformerBlock: the fundamental repeatable unit of every transformer model. Stack N of these blocks and you get a transformer. GPT-2 Small stacks 12. GPT-3 stacks 96. The block itself is always the same structure: two sub-layers, each wrapped with layer normalization and a residual connection.
What Does a TransformerBlock Do?
A TransformerBlock is the repeating unit inside every transformer. If you think of a transformer as a tall building, each TransformerBlock is one floor. Every floor has the same layout, the same two rooms, and the same wiring. Stack enough floors and you get GPT-2, GPT-3, LLaMA, or any other transformer-based model.
Each block contains exactly two sub-layers:
- Multi-Head Attention (MHA): Lets tokens communicate with each other. Every token looks at every other token (within the causal mask) and decides what information to gather. This is the "communication" step.
- Feed-Forward Network (FFN): Processes each token independently through an expand-activate-compress pipeline. This is the "computation" step, where each token digests the information it just gathered.
Both sub-layers are wrapped with two critical structural ingredients:
- Layer Normalization before each sub-layer (pre-norm architecture). This keeps activations in a stable range so training does not diverge.
- Residual connections around each sub-layer. The input to the sub-layer is added back to its output. This creates a "gradient highway" that allows deep stacking without vanishing gradients.
The high-level data flow looks like this:
Notice the input and output have exactly the same shape. This is essential. Because the shape is preserved, blocks can be stacked: the output of block 1 feeds directly into block 2, then into block 3, and so on. There is no adapter or reshape layer needed between blocks.
Think of it this way: the first few blocks learn surface-level patterns (syntax, common phrases). The middle blocks learn semantic relationships (what words mean in context). The deeper blocks learn abstract reasoning (logic, world knowledge). Each block incrementally refines the representation by adding a small correction via its two sub-layers.
The Class Structure
Here is the complete TransformerBlock class in C++ using LibTorch:
shared_ptr<MultiHeadAttention> MHA is the attention sub-layer. It is stored as a shared_ptr because LibTorch's module registration system requires shared ownership. When you call register_module("mha", MHA), the module system takes a copy of the shared pointer. Both the class member and the module registry now share ownership of the same MultiHeadAttention object. If we used a raw pointer or a unique_ptr, the registration would not work, because LibTorch expects to be able to hold a reference that keeps the module alive.
The MultiHeadAttention module itself contains multiple single attention heads, a projection layer, and runs the full Q-K-V attention computation. From the block's perspective, it is a black box: give it a tensor of shape (B, T, n_embd) and it returns one of the same shape.
shared_ptr<FeedForward> ff{nullptr} is the feed-forward sub-layer. Same shared_ptr pattern for the same reasons. The {nullptr} initializer means the pointer starts as null and gets assigned a real FeedForward object inside the constructor body. This two-step pattern (declare null, then construct) is common in LibTorch because module construction often depends on constructor parameters that are not available at declaration time.
The FeedForward module expands the n_embd-dimensional input to d_ff dimensions (typically 4x), applies a non-linearity (ReLU or GELU), then compresses back to n_embd. Each token is processed independently.
torch::nn::LayerNorm ln1{nullptr}, ln2{nullptr} are two separate layer normalization modules. They are declared as null and constructed in the body with LayerNormOptions({n_embd}).
Why two? Because ln1 normalizes the input before attention and ln2 normalizes the input before the feed-forward network. Each LayerNorm has its own learned gamma (scale) and beta (shift) parameters. They learn different normalization patterns because the two sub-layers process data differently. We will explore this in detail in the "Why Two Separate LayerNorms?" section below.
The argument {n_embd} tells LayerNorm which dimensions to normalize over. It specifies the normalized shape: the trailing dimensions of the input tensor across which the mean and variance are computed.
For a tensor of shape (B, T, n_embd), passing {n_embd} means: for each individual token (each of the B*T tokens), compute the mean and variance across its n_embd features, then normalize. Each token is normalized independently. The batch and sequence dimensions are untouched.
The curly braces {} are needed because LayerNormOptions expects a std::vector<int64_t>. Writing LayerNormOptions(n_embd) without braces would pass a plain int, which does not match the constructor signature. The braces create an initializer list that converts to the required vector type.
int num_heads, head_size, n_embd, d_ff stores the four hyperparameters as private members. These are initialized via the member initializer list: : num_heads(num_heads), head_size(head_size), n_embd(n_embd), d_ff(d_ff). This pattern binds constructor arguments to member variables before the constructor body executes.
Storing these values is useful for debugging, serialization, and any method that might need to reference the block's configuration after construction.
All four components are registered with register_module:
register_module("mha", MHA)- the attention sub-layerregister_module("ff", ff)- the feed-forward sub-layerregister_module("ln1", ln1)- first layer normregister_module("ln2", ln2)- second layer norm
Registration is what makes the module system work. Without it, calling block->parameters() would return an empty list. The optimizer would have nothing to optimize. Device transfers via block->to(torch::kCUDA) would miss unregistered sub-modules. Serialization via torch::save would skip them. Registration is not optional. It is the glue that connects your sub-modules to the LibTorch infrastructure.
Construction Order
The constructor body executes these steps in order:
- Create the MultiHeadAttention module with
make_sharedand assign to MHA. - Create the FeedForward module with
make_sharedand assign to ff. - Construct ln1 as a LayerNorm that normalizes over n_embd features.
- Construct ln2 as a separate LayerNorm with its own parameters.
- Register all four modules with string names.
After construction, the block is ready. Call forward(x) with any tensor of shape (B, T, n_embd) and it returns a tensor of the same shape.
Shape Trace Table
Let us trace every intermediate shape with concrete values: n_embd = 512, num_heads = 8, head_size = 64, d_ff = 2048, B = 4, T = 32.
| Step | Expression | Shape | What Happens |
|---|---|---|---|
| 0 | x (input) | (4, 32, 512) | Input tensor arrives at the block |
| 1 | ln1(x) | (4, 32, 512) | Normalize each token's 512 features. Mean and variance computed per token. |
| 2a | Inside MHA: Q, K, V per head | (4, 32, 64) each | Each of the 8 heads projects from 512 to head_size=64 |
| 2b | Inside MHA: attention scores | (4, 32, 32) per head | Q * K^T / sqrt(64) for each head. Token-to-token relevance. |
| 2c | Inside MHA: weighted values | (4, 32, 64) per head | softmax(scores) * V for each head |
| 2d | Inside MHA: concatenated | (4, 32, 512) | 8 heads of size 64 concatenated along last dim: 8*64 = 512 |
| 2e | MHA->forward(ln1(x)) | (4, 32, 512) | Output projection W_O maps 512 to 512 |
| 3 | x + MHA->forward(ln1(x)) | (4, 32, 512) | Residual add: element-wise addition with original x |
| 4 | ln2(x) (x is now x') | (4, 32, 512) | Normalize updated representation before FFN |
| 5a | Inside FFN: linear_1 | (4, 32, 2048) | Expand from 512 to 2048 (4x expansion) |
| 5b | Inside FFN: ReLU | (4, 32, 2048) | Zero out negative values. Sparsifies activations. |
| 5c | ff->forward(ln2(x)) | (4, 32, 512) | Compress from 2048 back to 512 |
| 6 | x + ff->forward(ln2(x)) | (4, 32, 512) | Second residual add with x' |
| 7 | return x (output) | (4, 32, 512) | Final output leaves the block. Same shape as input. |
Every intermediate tensor at the block level has shape (B, T, n_embd) = (4, 32, 512). The only place where shapes change is inside the sub-layers (the MHA's per-head projections and the FFN's expansion). But these internal shape changes are hidden from the block. From the block's perspective, both sub-layers are shape-preserving black boxes.
Expanded Shape View
Here is the same trace shown as a vertical flow, including the internal shapes that the sub-layers produce behind the scenes:
Parameter Count
Let us count every learnable parameter in a TransformerBlock with n_embd = 512, num_heads = 8, head_size = 64, d_ff = 2048.
Where Do the Parameters Live?
| Component | Parameters | Percentage |
|---|---|---|
| Multi-Head Attention | 1,049,088 | 33.3% |
| Feed-Forward Network | 2,099,712 | 66.6% |
| LayerNorm (both) | 2,048 | 0.065% |
| Total | 3,150,848 | 100% |
The FFN dominates at roughly two-thirds of the total parameters. This is because the FFN has a 4x expansion factor: the first linear layer maps from 512 to 2048 and the second maps back. Each of those linear layers has approximately 1 million parameters. The MHA, despite being the more conceptually complex component, has about half the parameters because the Q/K/V projections go to head_size (64), not the full d_ff (2048).
Parameter Efficiency Insight
The LayerNorm modules contribute only 2,048 parameters out of 3.15 million. That is less than 0.1%. Yet they are essential for training stability. Without LayerNorm, training a 12-layer model would require extremely careful learning rate tuning and warmup schedules. With LayerNorm, training "just works" for a wide range of hyperparameters. This is an extraordinary return on investment: 0.065% of parameters buys training stability for the other 99.935%.
Full Code
The complete TransformerBlock implementation with a main function for testing. This assumes SingleAttentionHead, MultiHeadAttention, and FeedForward are defined as in the previous blog posts.
(num_heads, head_size, n_embd, d_ff). The existing basic-blocks version uses (n_embd, num_heads, head_size, d_ff). Both are valid choices. What matters is consistency within your codebase. Pick one convention and stick with it.
What Comes Next
We now have a complete, self-contained, stackable TransformerBlock. It takes in (B, T, n_embd) and produces (B, T, n_embd). It includes multi-head attention for token communication, a feed-forward network for per-token computation, layer normalization for training stability, and residual connections for gradient flow.
The next step is to stack multiple TransformerBlocks into the full GPT model. That involves:
- Token embeddings: A lookup table that converts integer token IDs into dense vectors of dimension n_embd.
- Positional embeddings: Adding position information so the model knows the order of tokens (since attention is position-agnostic by itself).
- N stacked blocks: Creating a
std::vector<shared_ptr<TransformerBlock>>and running input through each block sequentially. - Final LayerNorm: A final normalization after the last block (pre-norm architectures typically add one more LayerNorm at the very end).
- Output head: A linear layer that maps from n_embd to the vocabulary size, producing logits for next-token prediction.
With those pieces in place, we will have a complete, trainable GPT model implemented from scratch in C++.