← All Posts

TransformerBlock in C++ (LibTorch)

We have built multi-head attention and the feed-forward network in previous posts. Now we combine them into the TransformerBlock: the fundamental repeatable unit of every transformer model. Stack N of these blocks and you get a transformer. GPT-2 Small stacks 12. GPT-3 stacks 96. The block itself is always the same structure: two sub-layers, each wrapped with layer normalization and a residual connection.

What Does a TransformerBlock Do?

A TransformerBlock is the repeating unit inside every transformer. If you think of a transformer as a tall building, each TransformerBlock is one floor. Every floor has the same layout, the same two rooms, and the same wiring. Stack enough floors and you get GPT-2, GPT-3, LLaMA, or any other transformer-based model.

Each block contains exactly two sub-layers:

  1. Multi-Head Attention (MHA): Lets tokens communicate with each other. Every token looks at every other token (within the causal mask) and decides what information to gather. This is the "communication" step.
  2. Feed-Forward Network (FFN): Processes each token independently through an expand-activate-compress pipeline. This is the "computation" step, where each token digests the information it just gathered.

Both sub-layers are wrapped with two critical structural ingredients:

  • Layer Normalization before each sub-layer (pre-norm architecture). This keeps activations in a stable range so training does not diverge.
  • Residual connections around each sub-layer. The input to the sub-layer is added back to its output. This creates a "gradient highway" that allows deep stacking without vanishing gradients.

The high-level data flow looks like this:

Input x [B, T, n_embd] LayerNorm (ln1) Multi-Head Attention residual + x + MHA(ln1(x)) LayerNorm (ln2) Feed-Forward residual + x' + FFN(ln2(x')) Output x'' [B, T, n_embd]

Notice the input and output have exactly the same shape. This is essential. Because the shape is preserved, blocks can be stacked: the output of block 1 feeds directly into block 2, then into block 3, and so on. There is no adapter or reshape layer needed between blocks.

How many blocks? GPT-2 Small uses 12 blocks. GPT-2 Medium uses 24. GPT-2 Large uses 36. GPT-3 uses 96. The number of blocks is one of the primary scaling knobs. More blocks mean more capacity for learning complex patterns, at the cost of more parameters and more compute.

Think of it this way: the first few blocks learn surface-level patterns (syntax, common phrases). The middle blocks learn semantic relationships (what words mean in context). The deeper blocks learn abstract reasoning (logic, world knowledge). Each block incrementally refines the representation by adding a small correction via its two sub-layers.

The Class Structure

Here is the complete TransformerBlock class in C++ using LibTorch:

class TransformerBlock : public torch::nn::Module{ shared_ptr<MultiHeadAttention> MHA; shared_ptr<FeedForward> ff{nullptr}; torch::nn::LayerNorm ln1{nullptr}, ln2{nullptr}; int num_heads, head_size, n_embd, d_ff; public: TransformerBlock(int num_heads, int head_size, int n_embd, int d_ff) : num_heads(num_heads), head_size(head_size), n_embd(n_embd), d_ff(d_ff){ MHA = make_shared<MultiHeadAttention>(num_heads, head_size, n_embd); ff = make_shared<FeedForward>(n_embd, d_ff); ln1 = torch::nn::LayerNorm( torch::nn::LayerNormOptions({n_embd}) ); ln2 = torch::nn::LayerNorm( torch::nn::LayerNormOptions({n_embd}) ); register_module("mha", MHA); register_module("ff", ff); register_module("ln1", ln1); register_module("ln2", ln2); } torch::Tensor forward(torch::Tensor x){ x = x + MHA->forward(ln1(x)); x = x + ff->forward(ln2(x)); return x; } };

shared_ptr<MultiHeadAttention> MHA is the attention sub-layer. It is stored as a shared_ptr because LibTorch's module registration system requires shared ownership. When you call register_module("mha", MHA), the module system takes a copy of the shared pointer. Both the class member and the module registry now share ownership of the same MultiHeadAttention object. If we used a raw pointer or a unique_ptr, the registration would not work, because LibTorch expects to be able to hold a reference that keeps the module alive.

The MultiHeadAttention module itself contains multiple single attention heads, a projection layer, and runs the full Q-K-V attention computation. From the block's perspective, it is a black box: give it a tensor of shape (B, T, n_embd) and it returns one of the same shape.

shared_ptr<FeedForward> ff{nullptr} is the feed-forward sub-layer. Same shared_ptr pattern for the same reasons. The {nullptr} initializer means the pointer starts as null and gets assigned a real FeedForward object inside the constructor body. This two-step pattern (declare null, then construct) is common in LibTorch because module construction often depends on constructor parameters that are not available at declaration time.

The FeedForward module expands the n_embd-dimensional input to d_ff dimensions (typically 4x), applies a non-linearity (ReLU or GELU), then compresses back to n_embd. Each token is processed independently.

torch::nn::LayerNorm ln1{nullptr}, ln2{nullptr} are two separate layer normalization modules. They are declared as null and constructed in the body with LayerNormOptions({n_embd}).

Why two? Because ln1 normalizes the input before attention and ln2 normalizes the input before the feed-forward network. Each LayerNorm has its own learned gamma (scale) and beta (shift) parameters. They learn different normalization patterns because the two sub-layers process data differently. We will explore this in detail in the "Why Two Separate LayerNorms?" section below.

The argument {n_embd} tells LayerNorm which dimensions to normalize over. It specifies the normalized shape: the trailing dimensions of the input tensor across which the mean and variance are computed.

For a tensor of shape (B, T, n_embd), passing {n_embd} means: for each individual token (each of the B*T tokens), compute the mean and variance across its n_embd features, then normalize. Each token is normalized independently. The batch and sequence dimensions are untouched.

The curly braces {} are needed because LayerNormOptions expects a std::vector<int64_t>. Writing LayerNormOptions(n_embd) without braces would pass a plain int, which does not match the constructor signature. The braces create an initializer list that converts to the required vector type.

int num_heads, head_size, n_embd, d_ff stores the four hyperparameters as private members. These are initialized via the member initializer list: : num_heads(num_heads), head_size(head_size), n_embd(n_embd), d_ff(d_ff). This pattern binds constructor arguments to member variables before the constructor body executes.

Storing these values is useful for debugging, serialization, and any method that might need to reference the block's configuration after construction.

All four components are registered with register_module:

  • register_module("mha", MHA) - the attention sub-layer
  • register_module("ff", ff) - the feed-forward sub-layer
  • register_module("ln1", ln1) - first layer norm
  • register_module("ln2", ln2) - second layer norm

Registration is what makes the module system work. Without it, calling block->parameters() would return an empty list. The optimizer would have nothing to optimize. Device transfers via block->to(torch::kCUDA) would miss unregistered sub-modules. Serialization via torch::save would skip them. Registration is not optional. It is the glue that connects your sub-modules to the LibTorch infrastructure.

Construction Order

The constructor body executes these steps in order:

  1. Create the MultiHeadAttention module with make_shared and assign to MHA.
  2. Create the FeedForward module with make_shared and assign to ff.
  3. Construct ln1 as a LayerNorm that normalizes over n_embd features.
  4. Construct ln2 as a separate LayerNorm with its own parameters.
  5. Register all four modules with string names.

After construction, the block is ready. Call forward(x) with any tensor of shape (B, T, n_embd) and it returns a tensor of the same shape.

Shape Trace Table

Let us trace every intermediate shape with concrete values: n_embd = 512, num_heads = 8, head_size = 64, d_ff = 2048, B = 4, T = 32.

StepExpressionShapeWhat Happens
0x (input)(4, 32, 512)Input tensor arrives at the block
1ln1(x)(4, 32, 512)Normalize each token's 512 features. Mean and variance computed per token.
2aInside MHA: Q, K, V per head(4, 32, 64) eachEach of the 8 heads projects from 512 to head_size=64
2bInside MHA: attention scores(4, 32, 32) per headQ * K^T / sqrt(64) for each head. Token-to-token relevance.
2cInside MHA: weighted values(4, 32, 64) per headsoftmax(scores) * V for each head
2dInside MHA: concatenated(4, 32, 512)8 heads of size 64 concatenated along last dim: 8*64 = 512
2eMHA->forward(ln1(x))(4, 32, 512)Output projection W_O maps 512 to 512
3x + MHA->forward(ln1(x))(4, 32, 512)Residual add: element-wise addition with original x
4ln2(x) (x is now x')(4, 32, 512)Normalize updated representation before FFN
5aInside FFN: linear_1(4, 32, 2048)Expand from 512 to 2048 (4x expansion)
5bInside FFN: ReLU(4, 32, 2048)Zero out negative values. Sparsifies activations.
5cff->forward(ln2(x))(4, 32, 512)Compress from 2048 back to 512
6x + ff->forward(ln2(x))(4, 32, 512)Second residual add with x'
7return x (output)(4, 32, 512)Final output leaves the block. Same shape as input.

Every intermediate tensor at the block level has shape (B, T, n_embd) = (4, 32, 512). The only place where shapes change is inside the sub-layers (the MHA's per-head projections and the FFN's expansion). But these internal shape changes are hidden from the block. From the block's perspective, both sub-layers are shape-preserving black boxes.

Why shape preservation matters: If block N outputs (B, T, n_embd) and block N+1 expects (B, T, n_embd), you can stack any number of blocks without adapters. This is why the TransformerBlock is called a "repeatable unit." You configure the architecture once and instantiate it N times.

Expanded Shape View

Here is the same trace shown as a vertical flow, including the internal shapes that the sub-layers produce behind the scenes:

Block input: (4, 32, 512) === Attention sub-layer === ln1(x): (4, 32, 512) normalize features Inside MHA: Per head Q/K/V: (4, 32, 64) project to head_size Attention scores: (4, 32, 32) token-to-token weights Head output: (4, 32, 64) weighted sum of values Concatenated (8 heads): (4, 32, 512) cat along last dim Projected: (4, 32, 512) W_O projection x + MHA output: (4, 32, 512) residual add === FFN sub-layer === ln2(x): (4, 32, 512) normalize features Inside FFN: After linear_1: (4, 32, 2048) expand 4x After ReLU: (4, 32, 2048) sparsify After linear_2: (4, 32, 512) compress back x + FFN output: (4, 32, 512) residual add Block output: (4, 32, 512)

Parameter Count

Let us count every learnable parameter in a TransformerBlock with n_embd = 512, num_heads = 8, head_size = 64, d_ff = 2048.

Multi-Head Attention (MHA): Per head: Q projection: 512 * 64 = 32,768 K projection: 512 * 64 = 32,768 V projection: 512 * 64 = 32,768 Total for 8 heads: 8 * 3 * 32,768 = 786,432 Output projection W_O: 512 * 512 + 512 = 262,656 MHA total: 1,049,088 Feed-Forward Network (FFN): linear_1: 512 * 2048 + 2048 = 1,050,624 linear_2: 2048 * 512 + 512 = 1,049,088 FFN total: 2,099,712 LayerNorm x 2: ln1: 512 (gamma) + 512 (beta) = 1,024 ln2: 512 (gamma) + 512 (beta) = 1,024 LN total: 2,048 ========================================================== Block total: 1,049,088 + 2,099,712 + 2,048 = 3,150,848 ==========================================================

Where Do the Parameters Live?

ComponentParametersPercentage
Multi-Head Attention1,049,08833.3%
Feed-Forward Network2,099,71266.6%
LayerNorm (both)2,0480.065%
Total3,150,848100%

The FFN dominates at roughly two-thirds of the total parameters. This is because the FFN has a 4x expansion factor: the first linear layer maps from 512 to 2048 and the second maps back. Each of those linear layers has approximately 1 million parameters. The MHA, despite being the more conceptually complex component, has about half the parameters because the Q/K/V projections go to head_size (64), not the full d_ff (2048).

Scaling to a full model: For a 12-block model (GPT-2 Small scale), the blocks alone contain 12 * 3,150,848 = 37,810,176 parameters. This does not include the token embedding table, the positional embedding table, or the final output head. Those add roughly another 40 million parameters, bringing the total to approximately 85 million (which aligns with GPT-2 Small's reported 117M after accounting for different d_ff configurations).

Parameter Efficiency Insight

The LayerNorm modules contribute only 2,048 parameters out of 3.15 million. That is less than 0.1%. Yet they are essential for training stability. Without LayerNorm, training a 12-layer model would require extremely careful learning rate tuning and warmup schedules. With LayerNorm, training "just works" for a wide range of hyperparameters. This is an extraordinary return on investment: 0.065% of parameters buys training stability for the other 99.935%.

Full Code

The complete TransformerBlock implementation with a main function for testing. This assumes SingleAttentionHead, MultiHeadAttention, and FeedForward are defined as in the previous blog posts.

#include <torch/torch.h> #include <vector> #include <memory> #include <string> #include <iostream> using namespace std; // Assumes SingleAttentionHead, MultiHeadAttention, and FeedForward // are defined as in the previous blog posts. class TransformerBlock : public torch::nn::Module{ shared_ptr<MultiHeadAttention> MHA; shared_ptr<FeedForward> ff{nullptr}; torch::nn::LayerNorm ln1{nullptr}, ln2{nullptr}; int num_heads, head_size, n_embd, d_ff; public: TransformerBlock(int num_heads, int head_size, int n_embd, int d_ff) : num_heads(num_heads), head_size(head_size), n_embd(n_embd), d_ff(d_ff){ MHA = make_shared<MultiHeadAttention>(num_heads, head_size, n_embd); ff = make_shared<FeedForward>(n_embd, d_ff); ln1 = torch::nn::LayerNorm( torch::nn::LayerNormOptions({n_embd}) ); ln2 = torch::nn::LayerNorm( torch::nn::LayerNormOptions({n_embd}) ); register_module("mha", MHA); register_module("ff", ff); register_module("ln1", ln1); register_module("ln2", ln2); } torch::Tensor forward(torch::Tensor x){ x = x + MHA->forward(ln1(x)); x = x + ff->forward(ln2(x)); return x; } }; // ---- Test harness ---- int main() { int n_embd = 512; int num_heads = 8; int head_size = n_embd / num_heads; // 64 int d_ff = 4 * n_embd; // 2048 auto block = make_shared<TransformerBlock>( num_heads, head_size, n_embd, d_ff); // Input: batch=4, seq_len=32, n_embd=512 torch::Tensor x = torch::randn({4, 32, n_embd}); torch::Tensor out = block->forward(x); cout << "Input shape: " << x.sizes() << endl; cout << "Output shape: " << out.sizes() << endl; // Input shape: [4, 32, 512] // Output shape: [4, 32, 512] // Count total parameters int64_t total = 0; for (auto& p : block->parameters()) { total += p.numel(); } cout << "Total parameters: " << total << endl; // Total parameters: 3150848 return 0; }
Note on constructor argument order: Our constructor takes (num_heads, head_size, n_embd, d_ff). The existing basic-blocks version uses (n_embd, num_heads, head_size, d_ff). Both are valid choices. What matters is consistency within your codebase. Pick one convention and stick with it.

What Comes Next

We now have a complete, self-contained, stackable TransformerBlock. It takes in (B, T, n_embd) and produces (B, T, n_embd). It includes multi-head attention for token communication, a feed-forward network for per-token computation, layer normalization for training stability, and residual connections for gradient flow.

The next step is to stack multiple TransformerBlocks into the full GPT model. That involves:

  • Token embeddings: A lookup table that converts integer token IDs into dense vectors of dimension n_embd.
  • Positional embeddings: Adding position information so the model knows the order of tokens (since attention is position-agnostic by itself).
  • N stacked blocks: Creating a std::vector<shared_ptr<TransformerBlock>> and running input through each block sequentially.
  • Final LayerNorm: A final normalization after the last block (pre-norm architectures typically add one more LayerNorm at the very end).
  • Output head: A linear layer that maps from n_embd to the vocabulary size, producing logits for next-token prediction.

With those pieces in place, we will have a complete, trainable GPT model implemented from scratch in C++.

Building blocks so far: Single Attention Head, Multi-Head Attention, Feed-Forward Network, and now TransformerBlock. Each layer built on top of the previous one. The GPT model is the final layer that combines embeddings + N blocks + output projection into one cohesive system.