FeedForward Network
This page walks through the C++ LibTorch implementation of the FeedForward network line by line. For the theory — what the FFN does, the expand-compress architecture, activation function choices (ReLU vs GELU vs SwiGLU), and position-wise independence — see the dedicated FeedForward Network theory blog.
1 · The Class Structure
Here is the complete FeedForward class. We will break it apart line by line
right after.
That is the entire implementation. Let us walk through every piece.
Class declaration and inheritance
We inherit from torch::nn::Module, which is the base class for all neural
network modules in LibTorch. This gives us parameter registration, serialization,
device management (.to(device)), and integration with the optimizer. Every
custom layer you build in LibTorch inherits from this base class.
The nullptr initialization pattern
This is a pattern you will see everywhere in LibTorch code. In LibTorch, module holders
like torch::nn::Linear are smart pointers (specifically,
torch::nn::ModuleHolder wrappers around std::shared_ptr).
Initializing them with nullptr means "this module exists as a member
variable, but it has not been constructed yet." We cannot construct them here in the
member declaration because we need constructor arguments (n_embd,
d_ff) that are only available inside the constructor body.
The alternative would be to use the constructor's member initializer list, but the
register_module approach in the constructor body is the idiomatic
LibTorch pattern. It registers the sub-module with the parent module's parameter
tracking system at the same time it constructs the layer.
Member variables
We store both dimensions as member variables. n_embd is the model's
embedding dimension (the width of the residual stream). d_ff is the
intermediate dimension of the FeedForward expansion. Keeping these around can be useful
for debugging, serialization, or reconstructing the module later.
The constructor
The constructor takes the two dimensions and uses a member initializer list to store them. Then, inside the body, we construct and register both linear layers.
Registering l1 (the expansion layer)
This does three things in one expression:
-
Constructs a
torch::nn::Linearlayer with input sizen_embdand output sized_ff. The weight matrix W1 has shape [d_ff, n_embd] and the bias vector b1 has shape [d_ff]. -
Registers the layer with the parent
Moduleunder the name"l1". This means the optimizer will find its parameters, serialization will include them, and.to(device)will move them to the correct device. -
Assigns the constructed and registered module holder back to the member
variable
l1.
Registering l2 (the compression layer)
Exactly the same pattern, but with reversed dimensions. l2 takes input of
size d_ff and produces output of size n_embd. The weight
matrix W2 has shape
[n_embd, d_ff] and the bias vector
b2 has shape
[n_embd].
.bias(false) to LinearOptions, so
LibTorch uses the default, which is bias = true. Some Transformer
implementations disable bias (GPT-style models often do), but including it is perfectly
valid. The bias adds a small number of extra parameters (d_ff + n_embd) but gives
the network slightly more flexibility.
l1 expands, l2 compresses
It is worth pausing to emphasize the asymmetry. The two layers are not identical. They have transposed roles:
| Layer | Input dim | Output dim | Role |
|---|---|---|---|
| l1 | n_embd | d_ff | Expansion (upscale) |
| l2 | d_ff | n_embd | Compression (downscale) |
If n_embd = 256 and d_ff = 1024, then l1 maps
from a 256-dimensional space into a 1024-dimensional space, and l2 maps
from that 1024-dimensional space back down to 256. The ReLU activation sits between
them, operating in the expanded 1024-dimensional space where the network has maximum
expressive power.
2 · The Forward Pass
The forward method is where the actual computation happens. Let us trace
it line by line with concrete tensor shapes. We will assume:
batch = 2(two sequences in the batch)S = 10(sequence length of 10 tokens)n_embd = 256(embedding dimension)d_ff = 1024(intermediate dimension, 4x expansion)
The input x has shape [batch, S, n_embd],
which in our example is [2, 10, 256]. This tensor
contains the attended output from the previous attention layer (after residual
addition and layer normalization, depending on where the TransformerBlock places those
operations).
We declare a local tensor variable to hold intermediate results. In LibTorch,
torch::Tensor is a lightweight handle (similar to a smart pointer), so
declaring it without initialization is cheap.
Line 1: The expansion
The first linear layer l1 applies the affine transformation
output = x · W1T + b1.
LibTorch's Linear module broadcasts across all leading dimensions
automatically. So even though x is 3D
([batch, S, n_embd]), the linear layer applies its
weight matrix to the last dimension. The output has shape
[batch, S, d_ff] = [2, 10, 1024].
What happened? Each of the 2 × 10 = 20 token vectors (each 256-dimensional) was independently projected into 1024 dimensions. The representation has been expanded. There are now 1024 "features" to work with instead of 256.
Line 2: ReLU activation
ReLU (Rectified Linear Unit) is the simplest nonlinear activation function:
Every element in the [2, 10, 1024] tensor is processed independently. Positive values are left unchanged. Negative values are set to zero. The shape does not change.
This is the critical nonlinearity. Without it, l2(l1(x)) would simplify
to a single matrix multiplication
W2 · W1 · x,
which is just another linear transformation. The ReLU between the two layers is what
gives the FeedForward its power. It creates a piecewise-linear function with
2d_ff possible linear regions, one for each subset of intermediate neurons
that are active (positive) vs. inactive (zeroed out).
In practice, roughly 50% of the intermediate neurons output zero for any given input (assuming pre-activation values are roughly symmetric around zero). This means the FeedForward naturally develops sparse activations: different inputs activate different subsets of the intermediate neurons. This sparsity has been shown to be important for the network's ability to store and retrieve factual knowledge.
Line 3: The compression
The second linear layer l2 projects the 1024-dimensional intermediate
representation back to the original 256 dimensions:
output = intermediate · W2T + b2.
Again, the linear layer broadcasts across the batch and sequence dimensions automatically.
The output shape is [2, 10, 256], which matches the input shape exactly. This is essential because the FeedForward output will be added back to the input via the residual connection. The dimensions must match for that addition to work.
Return
The transformed tensor is returned. In the TransformerBlock, this will be added to the residual stream: x = x + FFN(LayerNorm(x)).
d_ff is invisible from outside the module.
3 · Interactive Animation
The animation below shows the data flow for a single token through
the FeedForward network. We use n_embd = 4 and d_ff = 16
for visual clarity. Click Step to advance through each stage of the
computation. Click Reset to start over.
FeedForward Data Flow
Notice how in Step 3, several neurons turn gray. These are the neurons where the pre-ReLU value was negative. ReLU sets them to exactly zero, meaning they contribute nothing to the subsequent compression step. The connection lines from gray (zeroed) neurons to the output are faded, showing that only the active neurons influence the final result.
4 · Shape Trace Table
The following table tracks the tensor shape through every operation in the
forward method. We use the concrete example from Section 2:
batch = 2, S = 10, n_embd = 256,
d_ff = 1024.
| Operation | Code | Input Shape | Output Shape | Notes |
|---|---|---|---|---|
| Input | x | - | [2, 10, 256] | Attended output from MHA |
| Linear l1 | l1(x) | [2, 10, 256] | [2, 10, 1024] | W1: [1024, 256], b1: [1024] |
| ReLU | torch::relu(output) | [2, 10, 1024] | [2, 10, 1024] | Element-wise, shape unchanged |
| Linear l2 | l2(output) | [2, 10, 1024] | [2, 10, 256] | W2: [256, 1024], b2: [256] |
| Return | return output | - | [2, 10, 256] | Same shape as input |
torch::nn::Linear only
operates on the last dimension and broadcasts over everything else.
Shape trace with symbolic dimensions
For the general case, here is the same table with symbolic names:
| Operation | Shape |
|---|---|
| Input x | [B, S, n_embd] |
| After l1 | [B, S, d_ff] |
| After ReLU | [B, S, d_ff] |
| After l2 (output) | [B, S, n_embd] |
5 · Parameter Count
Let us count every learnable parameter in the FeedForward module. There are exactly two linear layers, each with a weight matrix and a bias vector.
Layer l1 (expansion)
| Parameter | Shape | Count |
|---|---|---|
| W1 (weight) | [d_ff, n_embd] | d_ff × n_embd |
| b1 (bias) | [d_ff] | d_ff |
Subtotal for l1: d_ff × n_embd + d_ff = d_ff × (n_embd + 1)
Layer l2 (compression)
| Parameter | Shape | Count |
|---|---|---|
| W2 (weight) | [n_embd, d_ff] | n_embd × d_ff |
| b2 (bias) | [n_embd] | n_embd |
Subtotal for l2: n_embd × d_ff + n_embd = n_embd × (d_ff + 1)
Total parameter count
Concrete example: n_embd = 256, d_ff = 1024
Counting it out
l1 weight: 1024 × 256 = 262,144
l1 bias: 1024
l2 weight: 256 × 1024 = 262,144
l2 bias: 256
Total: 262,144 + 1,024 + 262,144 + 256 = 525,568
That is approximately 525K parameters for a single FeedForward module. The bias terms (1,024 + 256 = 1,280) are a negligible fraction. The two weight matrices dominate, contributing 262,144 parameters each.
Comparison with attention
In a typical Transformer block with n_embd = 256 and 4 attention heads,
Multi-Head Attention has four projection matrices (Q, K, V, output), each of shape
[256, 256], totaling about 262K parameters (with biases). The FeedForward, at 525K,
has roughly twice the parameter count. This is the standard ratio:
the FFN contains about 2/3 of a Transformer block's total parameters.
n_embd, the FeedForward parameters
quadruple.
What if we disabled bias?
If we had passed .bias(false) to both LinearOptions, the
total would drop by d_ff + n_embd = 1024 + 256 = 1,280 parameters. For
our example, that is:
The difference is negligible. Whether to include bias is mostly a stylistic choice at
this scale. Some modern models (like LLaMA) omit biases throughout. Our implementation
keeps them because LibTorch defaults to bias = true.
6 · Full Code
Here is the complete, self-contained FeedForward class. Copy and paste
this into your project. It is ready to use with any LibTorch build.
Usage example
Integration with TransformerBlock (preview)
When we build the full TransformerBlock, the FeedForward will be used like this:
The ln1 and ln2 are LayerNorm modules. The
x = x + ... lines are the residual connections. Our FeedForward
module handles only the inner transformation. The surrounding residual and
normalization logic lives in the TransformerBlock.
7 · What Comes Next
We have now built three of the four core building blocks that make up a Transformer layer:
- Single-Head Attention - scaled dot-product attention with Q, K, V projections.
- Multi-Head Attention - multiple parallel attention heads with concatenation and output projection.
- FeedForward - the position-wise two-layer MLP with ReLU (this post).
The next step is to assemble the TransformerBlock. A TransformerBlock combines Multi-Head Attention and FeedForward with two crucial additional components:
- Residual connections that add the sub-layer input to its output, creating gradient highways that allow information and gradients to flow unimpeded through deep networks.
- Layer Normalization that stabilizes the hidden representations, preventing the magnitudes from exploding or vanishing across layers.
The TransformerBlock is where all the pieces click together. Once we have it, stacking N blocks gives us the full Transformer encoder or decoder. That is the next post in this series.