Transformer in C++ (LibTorch)
Introduction
This series builds a complete Transformer implementation in C++ using LibTorch, the C++ frontend for PyTorch. We start from a single attention head and work our way up to a full encoder/decoder stack, one building block at a time.
The goal is not to reinvent tensor math from scratch. Instead, we use LibTorch to handle the heavy lifting (tensor operations, autograd, GPU support) while we focus on understanding how each transformer component fits together in C++.
Why LibTorch
Writing raw matrix classes and manual softmax is educational but impractical for real transformer work. LibTorch gives us everything we need without leaving C++:
- torch::Tensor: GPU-ready N-dimensional arrays with broadcasting, slicing, and batched operations. No manual memory management for tensors.
- torch::nn::Module: Base class for neural network layers with automatic parameter tracking. Call
parameters()and get every learnable weight. - torch::nn::Linear: Pre-built linear projection (matrix multiply + optional bias). No need to implement W*x yourself.
- torch::softmax, torch::matmul, torch::tril: Numerically stable, optimized operations out of the box.
- Autograd: Automatic differentiation for training. Call
.backward()and gradients flow through your entire model.
Project Setup
Download LibTorch from pytorch.org (select C++ / LibTorch). Then set up a CMake project:
Build and run:
Series Roadmap
Each page builds on the previous one. By the GPT post, you have a working decoder-only language model; the final two posts generalize the same C++ building blocks into encoder-only and encoder-decoder Transformers.
1. SingleHeadAttention
Q, K, V projections, scaled dot-product, transpose(-2,-1) explained with animation, causal masking, and softmax.
Read →2. MultiHeadAttention
Running multiple heads in parallel, concatenating outputs, and the final linear projection.
Read →3. FeedForward
The position-wise expand-compress FFN: two linear layers with ReLU in between.
Read →4. TransformerBlock
Pre-norm architecture: LayerNorm + MHA + residual + LayerNorm + FFN + residual.
Read →5. PositionalEncoding
Sinusoidal encoding in LibTorch: arange, pow, sin/cos, and slice assignment.
Read →6. GPT (Full Model)
Token embedding + positional encoding + N transformer blocks + lm_head. The complete GPT.
Read →7. Encoder-Only Transformer
Bidirectional self-attention, padding masks, [CLS] pooling, classifier heads, and the full BERT-style forward pass.
Read →8. Encoder-Decoder Transformer
Source encoder memory, shifted target decoder inputs, masked self-attention, cross-attention, teacher forcing, and inference.
Read →