← All Posts
From Scratch · C++ · Transformers · Building Blocks

FeedForward Network

This page walks through the C++ LibTorch implementation of the FeedForward network line by line. For the theory — what the FFN does, the expand-compress architecture, activation function choices (ReLU vs GELU vs SwiGLU), and position-wise independence — see the dedicated FeedForward Network theory blog.

1 · The Class Structure

Here is the complete FeedForward class. We will break it apart line by line right after.

class FeedForward : public torch::nn::Module{ torch::nn::Linear l1{nullptr}, l2{nullptr}; int n_embd, d_ff; public: FeedForward(int n_embd, int d_ff) : n_embd(n_embd), d_ff(d_ff){ l1 = register_module( "l1", torch::nn::Linear( torch::nn::LinearOptions( n_embd, d_ff ) ) ); l2 = register_module( "l2", torch::nn::Linear( torch::nn::LinearOptions( d_ff, n_embd ) ) ); } torch::Tensor forward(torch::Tensor x){ torch::Tensor output; output = l1(x); output = torch::relu(output); output = l2(output); return output; } };

That is the entire implementation. Let us walk through every piece.

Class declaration and inheritance

class FeedForward : public torch::nn::Module{

We inherit from torch::nn::Module, which is the base class for all neural network modules in LibTorch. This gives us parameter registration, serialization, device management (.to(device)), and integration with the optimizer. Every custom layer you build in LibTorch inherits from this base class.

The nullptr initialization pattern

torch::nn::Linear l1{nullptr}, l2{nullptr};

This is a pattern you will see everywhere in LibTorch code. In LibTorch, module holders like torch::nn::Linear are smart pointers (specifically, torch::nn::ModuleHolder wrappers around std::shared_ptr). Initializing them with nullptr means "this module exists as a member variable, but it has not been constructed yet." We cannot construct them here in the member declaration because we need constructor arguments (n_embd, d_ff) that are only available inside the constructor body.

The alternative would be to use the constructor's member initializer list, but the register_module approach in the constructor body is the idiomatic LibTorch pattern. It registers the sub-module with the parent module's parameter tracking system at the same time it constructs the layer.

Member variables

int n_embd, d_ff;

We store both dimensions as member variables. n_embd is the model's embedding dimension (the width of the residual stream). d_ff is the intermediate dimension of the FeedForward expansion. Keeping these around can be useful for debugging, serialization, or reconstructing the module later.

The constructor

FeedForward(int n_embd, int d_ff) : n_embd(n_embd), d_ff(d_ff){

The constructor takes the two dimensions and uses a member initializer list to store them. Then, inside the body, we construct and register both linear layers.

Registering l1 (the expansion layer)

l1 = register_module( "l1", torch::nn::Linear( torch::nn::LinearOptions( n_embd, d_ff ) ) );

This does three things in one expression:

  1. Constructs a torch::nn::Linear layer with input size n_embd and output size d_ff. The weight matrix W1 has shape [d_ff, n_embd] and the bias vector b1 has shape [d_ff].
  2. Registers the layer with the parent Module under the name "l1". This means the optimizer will find its parameters, serialization will include them, and .to(device) will move them to the correct device.
  3. Assigns the constructed and registered module holder back to the member variable l1.

Registering l2 (the compression layer)

l2 = register_module( "l2", torch::nn::Linear( torch::nn::LinearOptions( d_ff, n_embd ) ) );

Exactly the same pattern, but with reversed dimensions. l2 takes input of size d_ff and produces output of size n_embd. The weight matrix W2 has shape [n_embd, d_ff] and the bias vector b2 has shape [n_embd].

Note on bias: This implementation includes bias terms for both linear layers. We did not pass .bias(false) to LinearOptions, so LibTorch uses the default, which is bias = true. Some Transformer implementations disable bias (GPT-style models often do), but including it is perfectly valid. The bias adds a small number of extra parameters (d_ff + n_embd) but gives the network slightly more flexibility.

l1 expands, l2 compresses

It is worth pausing to emphasize the asymmetry. The two layers are not identical. They have transposed roles:

LayerInput dimOutput dimRole
l1n_embdd_ffExpansion (upscale)
l2d_ffn_embdCompression (downscale)

If n_embd = 256 and d_ff = 1024, then l1 maps from a 256-dimensional space into a 1024-dimensional space, and l2 maps from that 1024-dimensional space back down to 256. The ReLU activation sits between them, operating in the expanded 1024-dimensional space where the network has maximum expressive power.

2 · The Forward Pass

The forward method is where the actual computation happens. Let us trace it line by line with concrete tensor shapes. We will assume:

torch::Tensor forward(torch::Tensor x){

The input x has shape [batch, S, n_embd], which in our example is [2, 10, 256]. This tensor contains the attended output from the previous attention layer (after residual addition and layer normalization, depending on where the TransformerBlock places those operations).

torch::Tensor output;

We declare a local tensor variable to hold intermediate results. In LibTorch, torch::Tensor is a lightweight handle (similar to a smart pointer), so declaring it without initialization is cheap.

Line 1: The expansion

output = l1(x); // [2, 10, 256] -> [2, 10, 1024]

The first linear layer l1 applies the affine transformation output = x · W1T + b1. LibTorch's Linear module broadcasts across all leading dimensions automatically. So even though x is 3D ([batch, S, n_embd]), the linear layer applies its weight matrix to the last dimension. The output has shape [batch, S, d_ff] = [2, 10, 1024].

What happened? Each of the 2 × 10 = 20 token vectors (each 256-dimensional) was independently projected into 1024 dimensions. The representation has been expanded. There are now 1024 "features" to work with instead of 256.

Line 2: ReLU activation

output = torch::relu(output); // [2, 10, 1024] -> [2, 10, 1024]

ReLU (Rectified Linear Unit) is the simplest nonlinear activation function:

ReLU(x) = max(0, x)

Every element in the [2, 10, 1024] tensor is processed independently. Positive values are left unchanged. Negative values are set to zero. The shape does not change.

This is the critical nonlinearity. Without it, l2(l1(x)) would simplify to a single matrix multiplication W2 · W1 · x, which is just another linear transformation. The ReLU between the two layers is what gives the FeedForward its power. It creates a piecewise-linear function with 2d_ff possible linear regions, one for each subset of intermediate neurons that are active (positive) vs. inactive (zeroed out).

In practice, roughly 50% of the intermediate neurons output zero for any given input (assuming pre-activation values are roughly symmetric around zero). This means the FeedForward naturally develops sparse activations: different inputs activate different subsets of the intermediate neurons. This sparsity has been shown to be important for the network's ability to store and retrieve factual knowledge.

Line 3: The compression

output = l2(output); // [2, 10, 1024] -> [2, 10, 256]

The second linear layer l2 projects the 1024-dimensional intermediate representation back to the original 256 dimensions: output = intermediate · W2T + b2. Again, the linear layer broadcasts across the batch and sequence dimensions automatically.

The output shape is [2, 10, 256], which matches the input shape exactly. This is essential because the FeedForward output will be added back to the input via the residual connection. The dimensions must match for that addition to work.

Return

return output; // [2, 10, 256]

The transformed tensor is returned. In the TransformerBlock, this will be added to the residual stream: x = x + FFN(LayerNorm(x)).

Shape summary: The input shape and output shape are always identical. The FeedForward is a shape-preserving transformation. It changes the content of the representation at each position, but not its dimensionality. The internal expansion to d_ff is invisible from outside the module.

3 · Interactive Animation

The animation below shows the data flow for a single token through the FeedForward network. We use n_embd = 4 and d_ff = 16 for visual clarity. Click Step to advance through each stage of the computation. Click Reset to start over.

FeedForward Data Flow

Input (n_embd=4) After l1 (d_ff=16) After ReLU (d_ff=16) Output (n_embd=4)
Ready. Click Step to begin.

Notice how in Step 3, several neurons turn gray. These are the neurons where the pre-ReLU value was negative. ReLU sets them to exactly zero, meaning they contribute nothing to the subsequent compression step. The connection lines from gray (zeroed) neurons to the output are faded, showing that only the active neurons influence the final result.

4 · Shape Trace Table

The following table tracks the tensor shape through every operation in the forward method. We use the concrete example from Section 2: batch = 2, S = 10, n_embd = 256, d_ff = 1024.

Operation Code Input Shape Output Shape Notes
Input x - [2, 10, 256] Attended output from MHA
Linear l1 l1(x) [2, 10, 256] [2, 10, 1024] W1: [1024, 256], b1: [1024]
ReLU torch::relu(output) [2, 10, 1024] [2, 10, 1024] Element-wise, shape unchanged
Linear l2 l2(output) [2, 10, 1024] [2, 10, 256] W2: [256, 1024], b2: [256]
Return return output - [2, 10, 256] Same shape as input
Key observation: The batch dimension and sequence dimension are completely untouched by the FeedForward. Only the last dimension (the feature dimension) changes: 256 → 1024 → 1024 → 256. The first two dimensions ride through unchanged because torch::nn::Linear only operates on the last dimension and broadcasts over everything else.

Shape trace with symbolic dimensions

For the general case, here is the same table with symbolic names:

Operation Shape
Input x [B, S, n_embd]
After l1 [B, S, d_ff]
After ReLU [B, S, d_ff]
After l2 (output) [B, S, n_embd]

5 · Parameter Count

Let us count every learnable parameter in the FeedForward module. There are exactly two linear layers, each with a weight matrix and a bias vector.

Layer l1 (expansion)

ParameterShapeCount
W1 (weight)[d_ff, n_embd]d_ff × n_embd
b1 (bias)[d_ff]d_ff

Subtotal for l1: d_ff × n_embd + d_ff = d_ff × (n_embd + 1)

Layer l2 (compression)

ParameterShapeCount
W2 (weight)[n_embd, d_ff]n_embd × d_ff
b2 (bias)[n_embd]n_embd

Subtotal for l2: n_embd × d_ff + n_embd = n_embd × (d_ff + 1)

Total parameter count

Total = (n_embd × d_ff + d_ff) + (d_ff × n_embd + n_embd) = 2 × n_embd × d_ff + d_ff + n_embd

Concrete example: n_embd = 256, d_ff = 1024

Counting it out

l1 weight: 1024 × 256 = 262,144
l1 bias: 1024
l2 weight: 256 × 1024 = 262,144
l2 bias: 256
Total: 262,144 + 1,024 + 262,144 + 256 = 525,568

That is approximately 525K parameters for a single FeedForward module. The bias terms (1,024 + 256 = 1,280) are a negligible fraction. The two weight matrices dominate, contributing 262,144 parameters each.

Comparison with attention

In a typical Transformer block with n_embd = 256 and 4 attention heads, Multi-Head Attention has four projection matrices (Q, K, V, output), each of shape [256, 256], totaling about 262K parameters (with biases). The FeedForward, at 525K, has roughly twice the parameter count. This is the standard ratio: the FFN contains about 2/3 of a Transformer block's total parameters.

Scaling rule of thumb: For the standard 4x expansion, the FeedForward parameter count is approximately 8 × n_embd2 (ignoring biases). If you double n_embd, the FeedForward parameters quadruple.

What if we disabled bias?

If we had passed .bias(false) to both LinearOptions, the total would drop by d_ff + n_embd = 1024 + 256 = 1,280 parameters. For our example, that is:

With bias: 525,568 parameters Without bias: 524,288 parameters Difference: 1,280 parameters (0.24%)

The difference is negligible. Whether to include bias is mostly a stylistic choice at this scale. Some modern models (like LLaMA) omit biases throughout. Our implementation keeps them because LibTorch defaults to bias = true.

6 · Full Code

Here is the complete, self-contained FeedForward class. Copy and paste this into your project. It is ready to use with any LibTorch build.

class FeedForward : public torch::nn::Module{ torch::nn::Linear l1{nullptr}, l2{nullptr}; int n_embd, d_ff; public: FeedForward(int n_embd, int d_ff) : n_embd(n_embd), d_ff(d_ff){ l1 = register_module( "l1", torch::nn::Linear( torch::nn::LinearOptions( n_embd, d_ff ) ) ); l2 = register_module( "l2", torch::nn::Linear( torch::nn::LinearOptions( d_ff, n_embd ) ) ); } torch::Tensor forward(torch::Tensor x){ torch::Tensor output; output = l1(x); output = torch::relu(output); output = l2(output); return output; } };

Usage example

int main(){ // Create a FeedForward module: n_embd=256, d_ff=1024 (4x expansion) FeedForward ffn(256, 1024); // Print parameter count int total = 0; for(const auto& p : ffn.parameters()) total += p.numel(); std::cout << "FeedForward parameters: " << total << std::endl; // Output: FeedForward parameters: 525568 // Create a random input: batch=2, seq_len=10, n_embd=256 auto x = torch::randn({2, 10, 256}); // Forward pass auto out = ffn.forward(x); std::cout << out.sizes() << std::endl; // Output: [2, 10, 256] return 0; }

Integration with TransformerBlock (preview)

When we build the full TransformerBlock, the FeedForward will be used like this:

// Inside TransformerBlock::forward() // Attention sub-layer with residual auto attn_out = mha.forward(ln1(x)); x = x + attn_out; // FeedForward sub-layer with residual auto ffn_out = ffn.forward(ln2(x)); x = x + ffn_out;

The ln1 and ln2 are LayerNorm modules. The x = x + ... lines are the residual connections. Our FeedForward module handles only the inner transformation. The surrounding residual and normalization logic lives in the TransformerBlock.

7 · What Comes Next

We have now built three of the four core building blocks that make up a Transformer layer:

The next step is to assemble the TransformerBlock. A TransformerBlock combines Multi-Head Attention and FeedForward with two crucial additional components:

The TransformerBlock is where all the pieces click together. Once we have it, stacking N blocks gives us the full Transformer encoder or decoder. That is the next post in this series.

Next in the series: TransformerBlock combines Multi-Head Attention + FeedForward with residual connections and LayerNorm. We will build it from scratch in C++ using the modules we have already implemented.