Compiler and Translation Units
Why the Compilation Model Matters
C++ has one of the most complex compilation models of any mainstream language. When you press "Build," the compiler does not simply read your source file and produce an executable. It performs a multi-stage pipeline: preprocessing, compilation, assembly, and linking. Each stage has its own rules, its own failure modes, and its own performance characteristics. Understanding these stages is essential for diagnosing build errors, structuring large projects, and writing code that compiles efficiently.
Unlike languages such as Java or Go where the compiler understands the module structure natively, C++ inherited its compilation model from C, which was designed in the early 1970s when memory was scarce and compilers were simple single-pass programs. The #include mechanism, translation units, and separate compilation are all consequences of that heritage. C++20 modules aim to fix this, but the traditional model will dominate codebases for years to come.
This article traces the journey of a .cpp file from source text to object file, explaining every stage in detail, and then explores the concept of translation units: what they are, why they matter, and how to manage them effectively.
Stage 1: The Preprocessor
Before the compiler ever sees your C++ code, the preprocessor runs. The preprocessor is essentially a text-manipulation tool. It does not understand C++ syntax, types, or semantics. It operates on tokens and performs textual substitution according to directives that begin with #.
The #include Directive
The most impactful preprocessor directive is #include. It does exactly one thing: it finds the named file and pastes its entire contents into the current file at the point of the directive. That is all. There is no intelligence, no deduplication, no dependency analysis. Just textual inclusion.
// main.cpp
#include <iostream> // pastes ~30,000 lines of standard library headers
#include "widget.h" // pastes the contents of widget.h
int main() {
Widget w;
std::cout << w.name() << "\n";
}
After preprocessing, main.cpp might expand from 7 lines to 40,000 lines. Every #include is recursive: if widget.h includes string.h, which includes stddef.h, all of those get pasted in too. This is why C++ compilation is famously slow: the compiler often processes tens of thousands of lines for even a trivial source file.
The angle-bracket form <iostream> searches system include paths. The quoted form "widget.h" searches the current directory first, then falls back to system paths. The exact search order is implementation-defined, but this is the common behavior on all major compilers.
The #define Directive and Macros
#define creates a macro: a name-to-text substitution rule. Macros come in two forms: object-like macros and function-like macros.
// Object-like macro: simple text replacement
#define MAX_BUFFER_SIZE 4096
#define PI 3.14159265358979
// Function-like macro: takes arguments
#define SQUARE(x) ((x) * (x))
#define MIN(a, b) ((a) < (b) ? (a) : (b))
// Usage:
int buf[MAX_BUFFER_SIZE]; // becomes: int buf[4096];
double area = PI * SQUARE(radius); // becomes: double area = 3.14... * ((radius) * (radius));
Function-like macros are dangerous because they perform textual substitution, not semantic evaluation. The extra parentheses in the definitions above are critical. Without them:
#define BAD_SQUARE(x) x * x
int result = BAD_SQUARE(2 + 3);
// Expands to: int result = 2 + 3 * 2 + 3; (= 11, not 25!)
Modern C++ strongly prefers constexpr variables and inline functions over macros. But macros remain essential for conditional compilation, include guards, and platform-specific code.
Conditional Compilation: #ifdef, #ifndef, #if
The preprocessor can include or exclude blocks of code based on whether a macro is defined or on the value of a constant expression:
#ifdef _WIN32
#include <windows.h>
using NativeHandle = HANDLE;
#elif defined(__linux__)
#include <unistd.h>
using NativeHandle = int;
#else
#error "Unsupported platform"
#endif
#if __cplusplus >= 202002L
// C++20 features available
#include <concepts>
#endif
#ifndef NDEBUG
#define LOG(msg) std::cerr << __FILE__ << ":" << __LINE__ << " " << msg << "\n"
#else
#define LOG(msg) ((void)0)
#endif
This is how cross-platform code works in C++. The preprocessor strips out the irrelevant platform blocks before the compiler ever sees them. It is a crude but effective form of compile-time configuration.
Other Preprocessor Directives
A few more directives deserve mention:
#undef NAMEremoves a macro definition. Useful when a library defines a macro that clashes with your code.#pragmais a compiler-specific directive. The most common is#pragma once(discussed below). Others include#pragma pack(structure packing),#pragma warning(MSVC warning control), and#pragma GCC diagnostic.#error "message"immediately stops compilation with the given message. Used to enforce configuration constraints.#line N "filename"overrides the current file and line number for error messages. Used by code generators.__FILE__,__LINE__,__func__,__DATE__,__TIME__are predefined macros that expand to the current file name, line number, function name, compilation date, and time respectively.
What Is a Translation Unit?
A translation unit (TU) is the fundamental unit of compilation in C++. It is defined as a source file after all preprocessing has been performed: all #include directives have been expanded, all macros have been substituted, and all conditional compilation blocks have been resolved. The result is one enormous stream of C++ tokens that the compiler processes as a single, self-contained unit.
// Before preprocessing: main.cpp is 10 lines
// After preprocessing: main.cpp might be 50,000 lines
// That entire expanded text IS the translation unit
Each .cpp file in your project produces exactly one translation unit. Header files (.h, .hpp) do not produce translation units on their own; they contribute to translation units when included by a .cpp file. A header included by five different .cpp files is processed five times, once in each translation unit.
The compiler processes each translation unit independently. When compiling main.cpp, the compiler has no knowledge of what util.cpp contains. This is the essence of separate compilation: each TU is compiled in isolation, and the linker later combines the resulting object files.
Translation Unit Boundaries and Visibility
Because each TU is compiled independently, names and definitions in one TU are invisible to another TU at compile time. If main.cpp wants to call a function defined in util.cpp, it must see a declaration (not a definition) of that function. This is typically provided by a header:
// util.h: declaration only
int compute(int x, int y);
// util.cpp: definition (in util.cpp's TU)
#include "util.h"
int compute(int x, int y) { return x * x + y * y; }
// main.cpp: uses the declaration to compile
#include "util.h"
int main() {
return compute(3, 4); // compiles fine: declaration is visible
// The actual function body is in a different TU, the linker resolves it
}
This is why header files exist: they provide the declarations that allow translation units to reference each other's definitions. The actual definitions live in .cpp files, compiled once, in one TU.
Stage 2: Compilation (Front-End + Optimizer)
After preprocessing produces the translation unit, the compiler front-end takes over. This is where actual C++ language processing happens. The front-end performs several sub-steps:
Lexical Analysis and Parsing
The lexer (tokenizer) breaks the stream of characters into tokens: keywords (class, int, return), identifiers (main, widget), literals (42, "hello"), operators (+, ->, ::), and punctuation ({, ;). The parser then arranges these tokens into an Abstract Syntax Tree (AST) according to the C++ grammar.
C++ is notoriously hard to parse. The grammar is context-sensitive (the meaning of T(x) depends on whether T is a type or a variable), and templates create additional parsing challenges (the "most vexing parse," angle-bracket ambiguity, etc.). This is partly why C++ compilers are slower than compilers for simpler languages.
Semantic Analysis
After parsing, the compiler performs semantic analysis: type checking, overload resolution, template instantiation, implicit conversions, and name lookup. This is where most compile errors originate. If you call a function with the wrong argument types, use an undeclared variable, or write an ill-formed template, the error is caught here.
Template instantiation is particularly expensive. When you write std::vector<int>, the compiler must generate a complete, specialized version of the vector class for int. If ten TUs each use std::vector<int>, the compiler instantiates it ten times (one per TU). The linker later discards duplicates, but the compilation work is not shared.
Intermediate Representation and Optimization
The front-end produces an Intermediate Representation (IR): a low-level, platform-independent representation of the code. GCC uses GIMPLE and RTL; Clang/LLVM uses LLVM IR. The optimizer then transforms the IR to improve performance: dead code elimination, constant folding, loop unrolling, inlining, vectorization, and many more passes.
Modern optimizers are incredibly sophisticated. At -O2, the compiler might perform 50-100 optimization passes. At -O3, it enables aggressive optimizations like auto-vectorization and function cloning. With LTO (Link-Time Optimization), the optimizer can see across translation unit boundaries, enabling interprocedural optimizations that would be impossible otherwise.
# See the IR produced by Clang:
clang++ -S -emit-llvm -O2 main.cpp -o main.ll
# See GCC's GIMPLE IR:
g++ -fdump-tree-optimized main.cpp
Stage 3: Assembly
After optimization, the compiler's back-end translates the IR into assembly language for the target architecture (x86-64, ARM, RISC-V, etc.). The assembly is a human-readable representation of machine instructions:
# x86-64 assembly for a simple function
_Z7computeii: # mangled name of compute(int, int)
imull %edi, %edi # edi = x * x
imull %esi, %esi # esi = y * y
addl %esi, %edi # edi = x*x + y*y
movl %edi, %eax # return value in eax
retq
You can inspect the assembly output of any compilation:
# Generate assembly:
g++ -S -O2 main.cpp -o main.s
clang++ -S -O2 main.cpp -o main.s
# Use Compiler Explorer (godbolt.org) for an interactive view
The assembler then takes this assembly text and produces an object file (.o on Unix, .obj on Windows). The object file contains machine code (binary instructions the CPU can execute), but it is not yet a complete executable. It contains unresolved references to functions and variables defined in other translation units. Resolving those references is the job of the linker, which we cover in the next article.
What Is Inside an Object File?
An object file (ELF format on Linux, COFF/PE on Windows, Mach-O on macOS) contains several sections:
.text- The machine code (compiled functions)..data- Initialized global/static variables..bss- Uninitialized global/static variables (takes no space on disk; the OS zeroes it at load time)..rodata- Read-only data (string literals, vtables, const globals).- Symbol table - A list of all symbols (functions, global variables) defined or referenced in this TU, along with their addresses (for defined symbols) or an "undefined" marker (for referenced symbols).
- Relocation table - Instructions telling the linker where to patch addresses once symbol locations are known.
# Inspect an object file's symbols:
nm -C main.o
# Inspect sections:
objdump -h main.o
readelf -S main.o # Linux ELF
dumpbin /headers main.obj # MSVC
▶ Compilation Pipeline: Step-by-Step
Step through the 4 phases of compiling a single .cpp file into an object file.
How #include Creates One Big TU
Consider a realistic scenario. You have a small project:
// config.h
#pragma once
constexpr int VERSION = 3;
// logger.h
#pragma once
#include <string>
void log(const std::string& msg);
// engine.h
#pragma once
#include "config.h"
#include "logger.h"
#include <vector>
class Engine {
std::vector<int> data_;
public:
void run();
};
// main.cpp
#include "engine.h"
#include <iostream>
int main() {
Engine e;
e.run();
}
After preprocessing, the TU for main.cpp includes: the entirety of <string> (which pulls in <cstring>, <memory>, <iterator>, etc.), the entirety of <vector>, the entirety of <iostream>, plus config.h, logger.h, and engine.h. A rough estimate: 50,000 to 80,000 lines of code, all to compile a 7-line source file.
This is the fundamental cost of the #include model. Every TU independently processes its full transitive closure of includes. Ten source files that all include <vector> will each independently parse and compile the vector header.
Header Guards and #pragma once
A header can be included multiple times in a single TU through transitive includes. Without protection, this causes redefinition errors:
// widget.h (NO guard)
class Widget { int x; };
// panel.h
#include "widget.h" // Widget defined here
// app.h
#include "widget.h" // Widget defined again!
#include "panel.h" // Widget defined a THIRD time!
The traditional solution is include guards (also called header guards):
// widget.h
#ifndef WIDGET_H
#define WIDGET_H
class Widget { int x; };
#endif // WIDGET_H
The first time widget.h is included, WIDGET_H is not defined, so the preprocessor enters the block, defines WIDGET_H, and processes the content. The second time, WIDGET_H is already defined, so the entire block is skipped. The macro name must be unique across the entire project; a common convention is PROJECT_MODULE_FILENAME_H.
The modern alternative is #pragma once:
// widget.h
#pragma once
class Widget { int x; };
#pragma once tells the preprocessor to include this file at most once per TU. It is not part of the C++ standard but is supported by every major compiler (GCC, Clang, MSVC, ICC). It is simpler and avoids the risk of macro name collisions. Some teams prefer include guards for strict standards compliance; others prefer #pragma once for simplicity.
#pragma once prevent multiple definitions within a single TU. They do NOT prevent a header from being processed in multiple TUs (that is by design: each TU needs the declarations). The redundant cross-TU processing is one reason C++ compile times are high.
Include Guards vs #pragma once: Trade-Offs
- Include guards are standard-conformant and work on every conforming preprocessor. They require a unique macro name and cannot be shared between files. They are also slightly slower in theory: the preprocessor must open the file, see the guard, and skip the content. In practice, compilers optimize this ("header guard detection").
- #pragma once is non-standard but universally supported. The preprocessor tracks which physical files have been included and skips re-opening them entirely. This can be faster, but it can fail with symlinks or complex build configurations where the same logical header has different filesystem paths.
In practice, either approach works. Most modern codebases use #pragma once. Google's style guide uses include guards. Both are acceptable.
Precompiled Headers (PCH)
Precompiled headers are a compiler extension that addresses the fundamental performance problem of the #include model. The idea is simple: compile a set of commonly included headers once, save the compiler's internal state (parsed AST, preprocessor state) to a binary file, and reuse that saved state for every TU that includes the same headers.
// pch.h: the precompiled header
#pragma once
#include <vector>
#include <string>
#include <map>
#include <memory>
#include <iostream>
#include <algorithm>
The compiler is instructed to precompile pch.h into a binary form (.pch on MSVC, .gch on GCC). Then every .cpp file that starts with #include "pch.h" loads the precompiled state instantly instead of re-parsing all those headers.
How to Use PCH on Major Compilers
# GCC: compile the header, then use it
g++ -x c++-header -O2 pch.h # creates pch.h.gch
g++ -O2 -include pch.h main.cpp # uses the precompiled header
# Clang: similar
clang++ -x c++-header -O2 pch.h -o pch.h.pch
clang++ -O2 -include-pch pch.h.pch main.cpp
# MSVC: create and use
cl /Yc"pch.h" pch.cpp # creates pch.pch
cl /Yu"pch.h" /Fp"pch.pch" main.cpp
PCH Rules and Limitations
- The PCH must be the first include in every source file that uses it. Any code or includes before the PCH include are errors (or silently ignored, depending on the compiler).
- The PCH must be compiled with the same compiler flags as the TUs that use it. Different optimization levels, different defines, or different target architectures invalidate the PCH.
- Only headers that rarely change should go in the PCH. If
pch.hincludes a frequently-modified project header, the PCH must be rebuilt every time that header changes, negating the benefit. - Standard library headers, third-party library headers, and stable internal headers are ideal PCH candidates.
In CMake, you can set up PCH easily:
target_precompile_headers(my_target PRIVATE
<vector>
<string>
<memory>
<iostream>
"my_stable_header.h"
)
PCH can reduce build times by 30-60% on large projects, especially those that heavily use the STL and Boost.
Seeing the Preprocessor Output
You can instruct the compiler to stop after preprocessing and dump the expanded TU:
# GCC/Clang: output preprocessed TU
g++ -E main.cpp -o main.i
clang++ -E main.cpp -o main.i
# MSVC:
cl /P main.cpp # produces main.i
# Count how many lines the preprocessor generated:
wc -l main.i # often 50,000+ lines for a simple file
Examining the .i file is the definitive way to understand what the compiler actually sees. It resolves all ambiguity about macro expansions, include ordering, and conditional compilation. When you have a mysterious compile error that seems impossible given your source code, check the preprocessed output.
Advanced Macro Techniques
The preprocessor supports a few advanced features that you will encounter in library code:
Stringification (#)
The # operator converts a macro argument to a string literal:
#define STRINGIFY(x) #x
#define TOSTRING(x) STRINGIFY(x)
const char* version = TOSTRING(VERSION);
// If VERSION is defined as 3, this becomes: const char* version = "3";
The double-macro pattern (TOSTRING calling STRINGIFY) is necessary because #x stringifies the literal argument text, not the expanded value. TOSTRING(VERSION) first expands VERSION to 3, then STRINGIFY(3) produces "3".
Token Pasting (##)
The ## operator concatenates two tokens into one:
#define CONCAT(a, b) a##b
#define MAKE_UNIQUE(prefix) CONCAT(prefix, __LINE__)
int MAKE_UNIQUE(temp_); // becomes temp_42 (if on line 42)
This is commonly used to generate unique variable names in macros and to build identifiers from prefixes and suffixes.
Variadic Macros
#define LOG_FMT(fmt, ...) fprintf(stderr, fmt "\n", __VA_ARGS__)
LOG_FMT("value = %d, name = %s", 42, "hello");
// Expands to: fprintf(stderr, "value = %d, name = %s" "\n", 42, "hello");
__VA_ARGS__ is the standard way to handle a variable number of arguments in a macro. C++20 adds __VA_OPT__ for cleaner handling of the zero-argument case.
Compilation Flags That Affect These Stages
The most relevant compiler flags for the preprocessing and compilation stages:
# Preprocessing control:
-D NAME=VALUE # define a macro from the command line
-U NAME # undefine a macro
-I /path # add include search path
-isystem /path # add system include path (suppresses warnings)
# Compilation control:
-O0 # no optimization (fastest compile, slowest code)
-O2 # standard optimization (good balance)
-O3 # aggressive optimization (may increase code size)
-Og # optimize for debug experience
-std=c++20 # select language standard
-Wall -Wextra # enable common warnings
-Werror # treat warnings as errors
-fPIC # position-independent code (required for shared libraries)
# Output control:
-c # compile only (produce .o, do not link)
-S # produce assembly (.s) instead of object file
-E # preprocess only (produce .i)
Designing Translation Units Well
How you organize code into translation units has real consequences for build times, link times, and code quality:
Keep Headers Lean
Every line in a header is processed in every TU that includes it. A 1,000-line header included by 100 source files means 100,000 lines of redundant processing. Use forward declarations instead of #include where possible:
// BAD: widget.h pulls in everything
#include <string>
#include <vector>
#include "engine.h"
class Widget {
std::string name_;
std::vector<int> data_;
Engine* engine_;
};
// BETTER: forward-declare what you can
#include <string>
#include <vector>
class Engine; // forward declaration: no need to include engine.h
class Widget {
std::string name_;
std::vector<int> data_;
Engine* engine_; // pointer/reference only: forward decl is sufficient
};
The PIMPL Idiom
The Pointer-to-Implementation (PIMPL) idiom hides all implementation details behind an opaque pointer, dramatically reducing header dependencies:
// widget.h
#pragma once
#include <memory>
class Widget {
public:
Widget();
~Widget();
void doWork();
private:
struct Impl;
std::unique_ptr<Impl> pimpl_;
};
// widget.cpp
#include "widget.h"
#include <vector>
#include <string>
#include "engine.h"
#include "database.h"
struct Widget::Impl {
std::vector<int> data;
std::string name;
Engine engine;
Database db;
};
Widget::Widget() : pimpl_(std::make_unique<Impl>()) {}
Widget::~Widget() = default;
void Widget::doWork() { /* uses pimpl_->data, pimpl_->engine, etc. */ }
Now any file that includes widget.h only needs <memory>. All the heavy headers (<vector>, <string>, engine.h, database.h) are confined to widget.cpp. Changes to those implementation headers only recompile widget.cpp, not every file that uses Widget.
Unity Builds
A unity build (or jumbo build) takes the opposite approach: instead of compiling each .cpp separately, it creates a single .cpp that includes all other source files:
// unity.cpp
#include "main.cpp"
#include "widget.cpp"
#include "engine.cpp"
#include "database.cpp"
This produces a single TU containing the entire program. Shared headers are parsed only once, template instantiations happen once, and the compiler can optimize across the entire codebase without LTO. The downside: any change to any source file requires recompiling everything, and name collisions between static/anonymous-namespace entities in different files become errors.
Unity builds are commonly used in game development (Unreal Engine uses them by default) and for CI/CD where a full rebuild from scratch is needed anyway.
Common Preprocessing and Compilation Errors
// "No such file or directory": header not found
#include "nonexistent.h"
// Fix: check include paths (-I flag) and file spelling
// "Redefinition of class X": missing header guard
// Fix: add #pragma once or include guards
// "Use of undeclared identifier 'foo'": declaration not visible
// Fix: include the correct header or add a forward declaration
// Thousands of errors from a single typo in a header
// This happens because the error cascades through everything that includes the header
// Fix: look at the FIRST error, not the last
Summary
- The preprocessor performs text substitution:
#includepastes files,#definecreates macros,#ifdefcontrols conditional compilation. - A translation unit is a
.cppfile after all preprocessing. Each TU is compiled independently into an object file. - The compilation stages are: preprocessing, lexing/parsing, semantic analysis, IR optimization, assembly generation, and object file creation.
- Header guards and
#pragma onceprevent duplicate definitions within a single TU. - Precompiled headers avoid redundant parsing of stable headers, reducing build times significantly.
- Good TU design (lean headers, forward declarations, PIMPL) reduces coupling and improves build times.
- The
-Eflag is your best friend for debugging preprocessor issues.