← All Posts

Transformer in C++ (LibTorch)

Introduction

This series builds a complete Transformer implementation in C++ using LibTorch, the C++ frontend for PyTorch. We start from a single attention head and work our way up to a full encoder/decoder stack, one building block at a time.

The goal is not to reinvent tensor math from scratch. Instead, we use LibTorch to handle the heavy lifting (tensor operations, autograd, GPU support) while we focus on understanding how each transformer component fits together in C++.

Why LibTorch

Writing raw matrix classes and manual softmax is educational but impractical for real transformer work. LibTorch gives us everything we need without leaving C++:

  • torch::Tensor: GPU-ready N-dimensional arrays with broadcasting, slicing, and batched operations. No manual memory management for tensors.
  • torch::nn::Module: Base class for neural network layers with automatic parameter tracking. Call parameters() and get every learnable weight.
  • torch::nn::Linear: Pre-built linear projection (matrix multiply + optional bias). No need to implement W*x yourself.
  • torch::softmax, torch::matmul, torch::tril: Numerically stable, optimized operations out of the box.
  • Autograd: Automatic differentiation for training. Call .backward() and gradients flow through your entire model.
Think of it this way: LibTorch is to C++ what PyTorch is to Python. Same tensor engine, same autograd, same model zoo. The API is nearly identical. If you can read PyTorch code, you can read LibTorch code.

Project Setup

Download LibTorch from pytorch.org (select C++ / LibTorch). Then set up a CMake project:

# CMakeLists.txt cmake_minimum_required(VERSION 3.18) project(transformer_cpp) find_package(Torch REQUIRED) set(CMAKE_CXX_STANDARD 17) add_executable(transformer main.cpp) target_link_libraries(transformer "${TORCH_LIBRARIES}")

Build and run:

mkdir build && cd build cmake .. -DCMAKE_PREFIX_PATH=/path/to/libtorch make ./transformer
Common issue: Make sure the LibTorch version matches your compiler. Use the C++17 ABI version if your compiler supports it. On Windows, use the Release version of LibTorch with MSVC.

Series Roadmap

Each page builds on the previous one. By the GPT post, you have a working decoder-only language model; the final two posts generalize the same C++ building blocks into encoder-only and encoder-decoder Transformers.

1. SingleHeadAttention

Q, K, V projections, scaled dot-product, transpose(-2,-1) explained with animation, causal masking, and softmax.

Read →

2. MultiHeadAttention

Running multiple heads in parallel, concatenating outputs, and the final linear projection.

Read →

3. FeedForward

The position-wise expand-compress FFN: two linear layers with ReLU in between.

Read →

4. TransformerBlock

Pre-norm architecture: LayerNorm + MHA + residual + LayerNorm + FFN + residual.

Read →

5. PositionalEncoding

Sinusoidal encoding in LibTorch: arange, pow, sin/cos, and slice assignment.

Read →

6. GPT (Full Model)

Token embedding + positional encoding + N transformer blocks + lm_head. The complete GPT.

Read →

7. Encoder-Only Transformer

Bidirectional self-attention, padding masks, [CLS] pooling, classifier heads, and the full BERT-style forward pass.

Read →

8. Encoder-Decoder Transformer

Source encoder memory, shifted target decoder inputs, masked self-attention, cross-attention, teacher forcing, and inference.

Read →