← All Posts

Positional Encoding in C++ (LibTorch)

This page walks through the C++ LibTorch implementation of positional encoding line by line. For the theory behind why transformers need positional encoding and how permutation invariance motivates it, see the dedicated Positional Encoding blog.

The Sinusoidal Formula

The original Attention Is All You Need paper encodes each position using a pair of equations — one for even dimensions, one for odd:

$$PE_{(pos,\, 2i)} = \sin\!\left(\frac{pos}{10000^{\,2i\,/\,d_{\text{model}}}}\right)$$ $$PE_{(pos,\, 2i+1)} = \cos\!\left(\frac{pos}{10000^{\,2i\,/\,d_{\text{model}}}}\right)$$

There are only three ingredients. Let's unpack each one.

1 — The variables

SymbolMeaningRange
$pos$ Token's position in the sequence $0,\; 1,\; 2,\; \ldots,\; \text{seq\_len} - 1$
$i$ Dimension pair index $0,\; 1,\; 2,\; \ldots,\; \tfrac{d_{\text{model}}}{2} - 1$
$d_{\text{model}}$ Embedding size (n_embd in our code) e.g. 256, 512, 768

Every pair index $i$ fills two slots: dimension $2i$ gets the sine value and dimension $2i+1$ gets the cosine value. So if $d_{\text{model}} = 512$, there are 256 pairs, producing all 512 values.

Why does $i$ only go to $\tfrac{d_{\text{model}}}{2} - 1$, and why the $-1$?
Because each value of $i$ covers two dimensions ($2i$ and $2i+1$), you only need half as many values of $i$ to fill all $d_{\text{model}}$ dimensions. The $-1$ is simply zero-based indexing: if $d_{\text{model}} = 512$, you need 256 pairs, so $i$ runs from $0$ to $255$, which is $\tfrac{512}{2} - 1 = 255$. It's the same reason an array of length $n$ has indices $0$ to $n-1$.

2 — The denominator (the key insight)

The denominator $10000^{2i/d_{\text{model}}}$ is what makes this formula powerful. It creates a geometric progression of frequencies across dimensions:

$$\text{divisor}_i = 10000^{\,2i\,/\,d_{\text{model}}}$$

As $i$ increases from $0$ to $\tfrac{d}{2}-1$, the divisor grows exponentially from $1$ to $10000$. A larger divisor means the sine/cosine oscillates slower as we move across positions. Remember: $i$ picks the frequency of the wave, but the wave itself oscillates over $pos$. So each dimension "watches" position changes at a different zoom level:

$i = 0$: divisor $= 1$. The formula becomes $\sin(pos)$, which completes a full cycle every $2\pi \approx 6$ positions. This means position 0 and position 3 already get very different values in this dimension. It's a high-frequency wave — sensitive to small position changes, so it helps the model tell apart tokens that are close together.

$i = d/4$: divisor $= 100$. The formula becomes $\sin(pos/100)$, which completes a cycle every $\approx 628$ positions. Positions 0 and 3 look almost identical here, but positions 0 and 300 look very different. This dimension is a low-frequency wave — it ignores local shuffles and only reacts to large jumps in position.

$i \to d/2$: divisor $\to 10000$. The formula becomes $\sin(pos/10000)$, cycling every $\approx 62{,}832$ positions. For any realistic sequence length, the value barely changes at all. This dimension is near-constant — it gives every position roughly the same value, acting like a shared "DC offset" that the model can use as a reference point.

The key connection: $i$ does not iterate over positions — it selects a frequency. But that frequency determines how fast the PE value changes as position varies. Low $i$ = rapid change across positions (fine-grained). High $i$ = slow change across positions (coarse-grained). The model sees all dimensions at once, so it gets both fine and coarse position signals for every token.

3 — Why sin and cos together?

Using both sin and cos for each frequency gives the model two orthogonal components. This matters because of a trigonometric identity:

$$\sin(pos + k) = \sin(pos)\cos(k) + \cos(pos)\sin(k)$$

For any fixed offset $k$, the encoding at position $pos + k$ is a linear combination of the encoding at position $pos$. This means the model can learn to attend to relative positions (e.g. "two tokens back") using a simple linear transformation — no complex distance computation needed.

Clock analogy: Think of each dimension pair as a clock hand spinning at a different speed. The seconds hand (low $i$) moves fast and repeats quickly. The hours hand (high $i$) moves slowly for long-range differentiation. Reading all the hands together gives a unique timestamp for every position.

The Code Line by Line

Here is the complete PositionalEncoding class. We will go through every single line afterward.

class PositionalEncoding : public torch::nn::Module{ /* input => batch_size x seq_len x n_embd output => batch_size x */ int n_embd; public: PositionalEncoding(int n_embd) : n_embd(n_embd){} torch::Tensor forward(torch::Tensor x){ int seq_len = x.size(1); torch::Tensor pos_indices = torch::arange(0, seq_len, 1).unsqueeze(1); // seq_len x 1 torch::Tensor dim_indices = torch::arange(0, n_embd, 2); // n_embd/2 torch::Tensor divisor = torch::pow(10000.0, dim_indices/(float)n_embd); // n_embd/2 torch::Tensor angles = pos_indices/divisor; torch::Tensor output = torch::zeros({seq_len, n_embd}, x.options()); output.slice(1, 0, n_embd, 2) = torch::sin(angles); output.slice(1, 1, n_embd, 2) = torch::cos(angles); return output; } };

Now let us dissect every line. Each step card below explains one operation.

class PositionalEncoding : public torch::nn::Module{

We inherit from torch::nn::Module. This is the standard base class for all neural network components in LibTorch. By inheriting from it, we can plug PositionalEncoding into a larger model graph, call named_parameters(), move it to GPU with to(), and use it with register_module() in parent modules.

However, as we will see later, this class has no learnable parameters. It could technically be a standalone function. Inheriting from nn::Module is a design choice for consistency, not a necessity.

int n_embd; PositionalEncoding(int n_embd) : n_embd(n_embd){}

The class has a single member: the embedding dimension n_embd, which must match your token embeddings since we add the positional encoding to them element-wise. The constructor stores it via the initializer list : n_embd(n_embd) and the body is empty — no layers to create, no weights to register, no buffers to allocate. Positional encoding is entirely computed on the fly.

int seq_len = x.size(1);

The input tensor x has shape (batch_size, seq_len, n_embd). Dimension 0 is the batch, dimension 1 is the sequence length, dimension 2 is the embedding. We extract seq_len because the positional encoding matrix must have one row per position. This makes the module flexible: it works with any sequence length at runtime.

torch::Tensor pos_indices = torch::arange(0, seq_len, 1).unsqueeze(1); // seq_len x 1

torch::arange(0, seq_len, 1) creates a 1D tensor [0, 1, 2, ..., seq_len-1] with shape (seq_len). These are the "pos" values from the formula.

.unsqueeze(1) adds a new dimension at index 1, turning the shape from (seq_len) to (seq_len, 1). This creates a column vector:

Before unsqueeze: [0, 1, 2, 3, 4, 5, 6, 7] shape: (8) After unsqueeze: [[0], shape: (8, 1) [1], [2], ... [7]]

The unsqueeze is critical for broadcasting in the division step later. Without it, you would get an element-wise division instead of the outer-product-style computation we need.

torch::Tensor dim_indices = torch::arange(0, n_embd, 2); // n_embd/2

torch::arange(0, n_embd, 2) creates [0, 2, 4, 6, ..., n_embd-2] with shape (n_embd/2). The step size of 2 gives us only the even indices. These correspond to the "2i" values in the formula.

Why only even indices? Because each pair of dimensions (2i and 2i+1) shares the same frequency. Dimension 0 and 1 share one frequency. Dimension 2 and 3 share another. We compute the angle once per pair, then apply sin to the even slot and cos to the odd slot.

If n_embd = 16, this produces: [0, 2, 4, 6, 8, 10, 12, 14] with 8 elements (16/2 = 8).

torch::Tensor divisor = torch::pow(10000.0, dim_indices/(float)n_embd); // n_embd/2

This computes the denominator 10000^(2i/d_model) from the formula. Let us trace it for n_embd = 16:

dim_indices = [0, 2, 4, 6, 8, 10, 12, 14] dim_indices / 16.0 = [0.0, 0.125, 0.25, 0.375, 0.5, 0.625, 0.75, 0.875] 10000 ^ (above) = [1.0, 5.62, 31.62, 177.8, 100.0, 562.3, 3162, 17783]

Notice the (float) cast on n_embd. Without it, dim_indices / n_embd would be integer division, truncating everything to 0. The cast forces floating-point division so we get the correct fractional exponents.

The divisors span a huge range: from 1.0 at the lowest dimension to nearly 10000 at the highest. This means low dimensions have small divisors (the angle changes rapidly with position) and high dimensions have large divisors (the angle changes slowly). This is the geometric progression of wavelengths that makes the encoding powerful.

torch::Tensor angles = pos_indices/divisor;

This is where broadcasting happens. pos_indices has shape (seq_len, 1) and divisor has shape (n_embd/2). PyTorch/LibTorch broadcasts the division to produce a matrix of shape (seq_len, n_embd/2).

Each cell (p, d) in this matrix contains: p / 10000^(2d/n_embd). This is exactly the angle argument to sin and cos in the formula. We will explain broadcasting in full detail in the next section.

torch::Tensor output = torch::zeros({seq_len, n_embd}, x.options());

Creates a zero-filled matrix of shape (seq_len, n_embd). We will fill in the columns: even columns get sin values, odd columns get cos values.

The x.options() argument is important. It copies the dtype and device from the input tensor x. If x lives on the GPU, the output tensor will also be on the GPU. If x is float32, the output will be float32. This ensures compatibility when we later add the positional encoding to the token embeddings.

output.slice(1, 0, n_embd, 2) = torch::sin(angles);

The .slice() call has four arguments. Here is what each one means:

ArgumentValueMeaning
dim1Operate along dimension 1 (columns). Dimension 0 would be rows.
start0Begin at column index 0.
endn_embdStop before column n_embd (i.e. go up to the last column).
step2Take every 2nd column: 0, 2, 4, 6, …

So .slice(1, 0, n_embd, 2) selects the even columns: 0, 2, 4, …, n_embd−2. This returns a view, not a copy — when we assign torch::sin(angles) to it, the sin values get written directly into those columns of output.

torch::sin(angles) has shape (seq_len, n_embd/2), which matches the slice exactly: seq_len rows × n_embd/2 selected columns. The shapes match, so the assignment works.

output.slice(1, 1, n_embd, 2) = torch::cos(angles);

Identical arguments except start = 1 instead of 0. This selects the odd columns: 1, 3, 5, 7, …, n_embd−1. The cosine of the same angles fills them.

After both assignments, the output matrix looks like this (schematically):

col: 0 1 2 3 4 5 ... sin(a0) cos(a0) sin(a1) cos(a1) sin(a2) cos(a2) ... sin(b0) cos(b0) sin(b1) cos(b1) sin(b2) cos(b2) ... ...

Where a0, a1, a2, ... are the angles for position 0 at different frequencies, and b0, b1, ... are for position 1, and so on.

return output;

The returned tensor has shape (seq_len, n_embd). Note that this is 2D, not 3D. There is no batch dimension. That is intentional: the positional encoding is the same for every sample in the batch. When the caller adds it to the token embeddings (shape: batch_size x seq_len x n_embd), PyTorch broadcasts the 2D encoding across the batch dimension automatically.

Broadcasting Explained

The most subtle line in the entire class is the division pos_indices / divisor. It relies on broadcasting, and if you have not seen this before, the shapes might look incompatible. Let us walk through it in full detail.

The shapes involved

TensorShapeContents (example, seq_len=4, n_embd=8)
pos_indices(4, 1)[[0], [1], [2], [3]]
divisor(4)[1.0, 5.62, 31.62, 177.8]

Broadcasting rules

When two tensors have different numbers of dimensions, PyTorch aligns them from the right and pads with 1 on the left. For our case:

pos_indices: (4, 1) divisor: (4) => pad to (1, 4) Now compare dimensions right to left: dim 1: pos_indices has 1, divisor has 4 => broadcast pos_indices to 4 dim 0: pos_indices has 4, divisor has 1 => broadcast divisor to 4 Result shape: (4, 4)

What actually happens

Broadcasting "stretches" each tensor along dimensions where its size is 1. The position column [[0],[1],[2],[3]] gets copied across 4 columns. The divisor row [1.0, 5.62, 31.62, 177.8] gets copied across 4 rows. Then element-wise division happens:

divisor: [ 1.0 5.62 31.62 177.8 ] pos = 0: [ 0/1 0/5.6 0/31.6 0/178 ] = [ 0.000 0.000 0.000 0.000 ] pos = 1: [ 1/1 1/5.6 1/31.6 1/178 ] = [ 1.000 0.178 0.032 0.006 ] pos = 2: [ 2/1 2/5.6 2/31.6 2/178 ] = [ 2.000 0.356 0.063 0.011 ] pos = 3: [ 3/1 3/5.6 3/31.6 3/178 ] = [ 3.000 0.534 0.095 0.017 ]

Look at the pattern: the first column changes rapidly (0, 1, 2, 3), while the last column barely changes (0.000, 0.006, 0.011, 0.017). This is the frequency difference in action. After applying sin and cos, the first column oscillates fast and the last column oscillates slowly.

No memory duplication: Broadcasting does not actually copy data in memory. LibTorch uses stride tricks to make the tensor "appear" expanded without allocating new storage. This makes broadcasting both memory efficient and fast.

Interactive Heatmap Animation

The heatmap below shows the actual positional encoding values for 8 positions and 16 dimensions. Each cell's color represents the encoded value: blue for -1, white for 0, red for +1. Use the buttons to step through positions and watch how the pattern changes.

Position: none highlighted

What to notice in the heatmap

  • Leftmost columns (low dimensions) oscillate rapidly. The colors alternate between red and blue across consecutive positions. These are the "fast" frequencies.
  • Rightmost columns (high dimensions) change very slowly. They stay nearly white or light blue across all 8 positions. These are the "slow" frequencies.
  • Position 0 is special: all sin values are 0 (white) and all cos values are 1 (red).
  • Each row is unique. No two positions share the same color pattern. This is the "fingerprint" property.
  • Even columns (sin) and odd columns (cos) are paired. Column 0 (sin) and column 1 (cos) share the same frequency but are 90 degrees out of phase.

Frequency Intuition

The best analogy for positional encoding is a binary counter. Think about how binary numbers represent position:

Position Bit3 Bit2 Bit1 Bit0 0 0 0 0 0 1 0 0 0 1 2 0 0 1 0 3 0 0 1 1 4 0 1 0 0 5 0 1 0 1 6 0 1 1 0 7 0 1 1 1

Bit 0 (the lowest bit) flips every step. Bit 1 flips every 2 steps. Bit 2 flips every 4 steps. Bit 3 flips every 8 steps. Together, they uniquely identify every position from 0 to 7.

Sinusoidal positional encoding works the same way, but using smooth waves instead of binary flips:

  • Low dimensions (dim 0, 1) are like Bit 0: they oscillate rapidly, changing noticeably between every pair of adjacent positions. Period of approximately 2π (about 6.3 positions).
  • Mid dimensions (dim d/2) are like Bit 2: they oscillate more slowly, changing noticeably only over many positions. Period of approximately 200π.
  • High dimensions (dim d-2, d-1) are like the highest bit: they barely change at all across a typical sequence. Period approaching 20000π.
Why a geometric progression? Using exponentially spaced frequencies (1, 5.6, 31.6, 177.8, 1000, ...) instead of linearly spaced ones (1, 2, 3, 4, ...) gives each dimension a distinct "resolution." Linear spacing would make adjacent dimensions nearly redundant. Geometric spacing ensures that each dimension captures a fundamentally different scale of position.
Three frequencies from positional encoding dim 0 (fast) dim 4 (medium) dim 14 (slow) position (0 to 30)

The fast wave (dim 0) completes nearly 5 full cycles across 30 positions. The slow wave (dim 14) barely curves at all. This multi-scale representation is what allows the transformer to detect both fine-grained (adjacent token) and coarse-grained (distant token) positional relationships.

Uniqueness guarantee

Because the frequencies form a geometric progression with an irrational base (10000 raised to various rational powers), the combined encoding is unique for every integer position. No two positions will ever produce the exact same vector. This is similar to how a set of incommensurate sine waves produces a quasi-periodic signal that never exactly repeats.

Extrapolation to longer sequences

Because sinusoidal encoding is a mathematical formula (not a learned lookup table), it can generate encodings for positions it never saw during training. If you train on sequences of length 512 and then test on length 1024, the encoding still works. The sin/cos values are well-defined for any non-negative integer position. This is a practical advantage over learned positional embeddings.

The cost of slow waves

At high dimensions (large $i$), the divisor approaches 10000 and the wave barely moves. For a 512-token sequence, $\sin(\text{pos}/10000)$ stays near 0 and $\cos(\text{pos}/10000)$ stays near 1 at every position. Those dimensions produce almost the same value for position 0 and position 511.

The full PE vector is still unique per position — but uniqueness alone doesn't help. Here's why.

Attention decides how much token A should attend to token B by computing the dot product $Q_A \cdot K_B$. That dot product is a sum across all dimensions:

$$Q_A \cdot K_B \;=\; \sum_{d=0}^{d_{\text{model}}-1} Q_A[d] \;\times\; K_B[d]$$

Each dimension contributes one term to this sum. Now, $Q$ and $K$ are linear projections of (token embedding + PE). In the low-$i$ dimensions, the PE values are very different for different positions — so the terms in the sum carry a strong positional signal. In the high-$i$ dimensions, the PE values are nearly identical across all positions — so those terms contribute roughly the same value regardless of whether token B is at position 5 or position 500. They add a near-constant offset to every dot product, making it harder for the model to distinguish "nearby" from "far away" based on position alone.

Suppose $d_{\text{model}} = 512$. Dimensions 0–200 have fast-varying PE and produce dot-product terms that differ by, say, $\pm 0.5$ depending on position distance. Dimensions 400–511 have near-constant PE and produce terms that differ by only $\pm 0.001$. The total dot product sums all 512 terms — the 112 near-constant terms contribute noise that the 200 useful terms must overcome. The positional signal gets diluted inside the sum.

That said, this is a minor issue in practice. The token embedding (which PE is added to) still carries rich semantic content in every dimension. And the learned Q/K/V projection weights can downweight dimensions they find uninformative. The model routes around the problem.

Modern alternatives solve bigger problems:
  • RoPE — encodes relative distance directly in the Q·K dot product, not absolute position
  • ALiBi — adds a simple linear distance bias to attention scores, no PE vectors needed
  • Learned PE — lets the model discover its own positional pattern, but cannot extrapolate beyond training length
These replaced sinusoidal PE mainly for better relative-position encoding and length generalization, not because of wasted dimensions.

Why long sequences are hard (and it's not PE's fault)

You might wonder: if sinusoidal PE works for any position, why are long-context transformers (say, 1M tokens) so much harder to train than short ones (512 tokens)? The formula scales fine — the real bottlenecks lie elsewhere:

Standard self-attention computes an $(S \times S)$ score matrix. At $S = 512$, that is 262K entries. At $S = 1{,}000{,}000$, it is 1 trillion. Memory and compute explode quadratically. This is the primary reason long sequences are hard — and why Flash Attention, Ring Attention, and Sparse Attention exist.

Softmax distributes probability across all positions. With 1M tokens, each position's average attention weight is ~$10^{-6}$. The model struggles to concentrate on the few tokens that actually matter. It is like searching a library by giving equal consideration to every book.

Longer sequences mean longer dependency chains for gradients to traverse. The loss landscape becomes harder to navigate — vanishing or exploding gradients across 1M steps of backpropagation, even with residual connections and layer normalization.

Each sample is 1M tokens, so fewer samples fit per batch — leading to noisier gradient estimates and slower convergence. You also need training data with meaningful long-range dependencies, which is scarce.

Where PE fits in: Sinusoidal PE actually scales fine to any length (the formula works for any position). The "wasted dimensions" problem we discussed above even improves at longer lengths, since the slow waves finally get room to oscillate. The real PE challenge for long sequences is that learned positional embeddings (used by GPT-2, BERT) have a fixed lookup table and literally cannot extrapolate beyond their trained length. That is why RoPE and ALiBi became essential for long-context models — they handle position without a length ceiling.

Why No Learnable Parameters?

If you look at the class, there is no call to register_module(), no register_parameter(), no register_buffer(). The class inherits from torch::nn::Module but registers nothing. Let us think about what that means.

No learnable parameters

When you call model->parameters() on a module, it returns all registered parameters. These are the tensors that the optimizer updates during training. Our PositionalEncoding has zero such tensors. The sin/cos values are computed fresh every forward pass from a deterministic formula. There is nothing to learn.

// If we inspect the parameters: auto pe = std::make_shared<PositionalEncoding>(512); for (auto& p : pe->named_parameters()) { cout << p.key() << endl; } // Output: (nothing - no parameters)

Comparison with learned positional embeddings

Some models (like GPT-2) use a learned positional embedding instead. That looks like:

// Learned approach: torch::nn::Embedding pos_emb = register_module("pos_emb", torch::nn::Embedding(1024, 512)); // max_seq_len x n_embd // This creates a 1024 x 512 parameter matrix // The optimizer learns the best position vectors during training

The learned approach creates a parameter matrix with max_seq_len * n_embd trainable values. The sinusoidal approach has zero trainable values. The tradeoffs are:

PropertySinusoidalLearned
Parameters0max_seq_len * n_embd
ExtrapolationWorks for any lengthLimited to max_seq_len
FlexibilityFixed formulaAdapts to data
Training costNo gradient neededGradient updates each step
Could this be a plain function? Yes. Since there are no parameters, no buffers, and no sub-modules, PositionalEncoding could simply be a free function: torch::Tensor positional_encoding(torch::Tensor x, int n_embd). Wrapping it in a class is a stylistic choice. It keeps the interface consistent with other components (all are nn::Module subclasses), makes it easy to plug into a register_module() chain, and groups the logic with its configuration (n_embd).

Why inherit from nn::Module then?

  • Consistency: Every other component in the transformer (attention, FFN, layer norm) is a Module. Keeping positional encoding as a Module means the parent module can register it uniformly.
  • Device tracking: When you call model->to(torch::kCUDA), it recursively moves all sub-modules. If PositionalEncoding were a free function, you would need to manually handle device placement of the output tensor (which we already do via x.options()).
  • Serialization: If you later wanted to cache the positional encoding as a buffer (using register_buffer), the Module infrastructure is already in place.
  • Clarity: When someone reads the GPT model definition and sees register_module("pe", ...), they immediately know it is a component, even if it has no weights.

Shape Trace Table

The table below traces every tensor's shape through the forward pass. We use concrete values: batch_size=4, seq_len=32, n_embd=512.

LineVariableShapeNotes
Input x (4, 32, 512) batch_size x seq_len x n_embd. From token embeddings.
1 seq_len scalar: 32 Extracted from x.size(1)
2 arange(0, 32, 1) (32) [0, 1, 2, ..., 31]
3 pos_indices (32, 1) After .unsqueeze(1). Column vector.
4 dim_indices (256) [0, 2, 4, ..., 510]. n_embd/2 = 256 elements.
5 dim_indices/(float)n_embd (256) [0/512, 2/512, 4/512, ..., 510/512]
6 divisor (256) 10000 raised to each exponent. Range: 1.0 to ~9770.
7 angles (32, 256) Broadcasting: (32,1) / (256) = (32,256). Each cell = pos/divisor.
8 output (32, 512) Zero-initialized. Will be filled with sin/cos values.
9 output.slice(1,0,512,2) (32, 256) View of even columns. 256 columns selected.
10 torch::sin(angles) (32, 256) Matches the slice shape. Assigned to even columns.
11 output.slice(1,1,512,2) (32, 256) View of odd columns. 256 columns selected.
12 torch::cos(angles) (32, 256) Matches the slice shape. Assigned to odd columns.
Output return output (32, 512) 2D: seq_len x n_embd. No batch dimension.
Note the output is 2D, not 3D. The input is (batch, seq_len, n_embd) but the output is (seq_len, n_embd). The batch dimension is absent because every sample in the batch gets the same positional encoding. When the caller does x + pe.forward(x), PyTorch broadcasts the 2D tensor across the batch dimension.

Memory cost

The only significant allocation is the output tensor: seq_len * n_embd floats. For seq_len=32 and n_embd=512, that is 32 * 512 * 4 bytes = 64 KB. Trivial. Even for seq_len=2048 and n_embd=768 (GPT-2 scale), it is only 2048 * 768 * 4 = 6 MB. Positional encoding is never the memory bottleneck.

What Comes Next

We now have positional encoding, the piece that injects sequence order into the transformer. Combined with the other building blocks we have built, we are getting close to a full model. The final step is the GPT class, which assembles everything:

  • Token embedding layer: Converts integer token IDs into dense vectors of size n_embd.
  • Positional encoding: The module we just built. Added to the token embeddings.
  • Stack of transformer blocks: Each block contains multi-head attention, feed-forward network, layer normalization, and residual connections.
  • Final linear head: Projects from n_embd back to vocabulary size for next-token prediction.

The full GPT model is where all these pieces come together into a single forward pass. That is the next post in this series.

The pipeline so far: Token IDs enter the embedding layer, get positional encoding added, pass through N transformer blocks (each with attention + FFN), and finally go through a linear layer to produce logits over the vocabulary. We have now built every individual piece. Time to assemble.