← All Posts
Deep Learning · Transformers· Training

Tokenization: How Text Becomes the Sequence a Model Learns

A language model predicts token IDs. Tokenization chooses the discrete units that become its sequence. It affects sequence length, the vocabulary head, context usage, and the interpretation of likelihood.

Separate text processing from learned representations

A typical pipeline normalizes or preprocesses text, splits it according to a tokenizer, maps pieces to integer IDs, and then looks up learned embedding rows. The tokenizer determines the IDs. The embedding table turns them into vectors. Updating model embeddings during training does not automatically change the tokenizer's segmentation rules.

Consider “redoing.” A tokenizer might produce one token, character tokens, or pieces such as “re”, “do”, and “ing.” Those are possible segmentations, not claims about a particular tokenizer. The model sees whichever ID sequence its fixed tokenizer actually emits. Changing that tokenizer after training requires a compatible vocabulary mapping and usually further training.

Train a tiny byte-pair encoding vocabulary

Take a deliberately simple corpus containing low three times and lower twice. Start with characters plus an end marker: l o w </w> and l o w e r </w>. Count adjacent pairs with word frequencies. Both l o and o w occur five times; specify a tie-break rule and choose l o.

Merge that pair into lo. Now lo w occurs five times, so merge it into low. At this stage the words are segmented as low </w> and low e r </w>. Subsequent merges depend on the remaining pair counts and tie rules. The training algorithm creates an ordered merge list, not a dictionary that always greedily selects the longest visible substring.

During tokenization, apply the learned merge priorities within the tokenizer's permitted regions. Modern byte-level BPE systems may use bytes as the initial alphabet and have explicit pre-tokenization rules. The character/end-marker example exposes the merge idea without pretending to reproduce those implementation details. See the subword BPE paper.

Try the mechanism Apply the learned merges in order

The resulting pieces map to categorical IDs, which index the learned embedding table. Tokenizer rules and embeddings are separate objects. Enable JavaScript to change the example inputs; the complete calculation remains in the article.

BPE is not every subword algorithm

A unigram tokenizer assigns probabilities to candidate pieces and chooses or samples a segmentation according to their sequence probability. This differs from applying a ranked list of pair merges. Subword Regularization describes a unigram language-model approach and sampling alternate segmentations during training.

WordPiece, byte-level BPE, and unigram tokenization can produce superficially similar pieces while having different training and inference algorithms. Also separate the segmentation algorithm from its software container: a tokenizer library can support multiple algorithms.

A larger vocabulary moves costs around

Let $V$ be vocabulary size and $d$ the embedding width. The embedding table stores $Vd$ parameters. A full vocabulary projection for each position has a cost proportional to $Vd$. A larger vocabulary can encode frequent strings in fewer tokens, reducing the sequence length seen by the transformer, but expands the embedding and output machinery.

For a fixed document, reducing its token count from 1000 to 800 reduces a dense attention score matrix from one million to 640,000 entries, a 36% reduction. That is an arithmetic example for the same underlying text; it does not prove a larger vocabulary yields that compression or improves quality. The vocabulary head and data coverage still matter.

Rare spellings, code, whitespace patterns, and languages with different writing systems can have very different tokens-per-character ratios. A context limit measured in tokens is therefore not a fixed number of words or pages.

Perplexities require compatible units

Average token loss is $\mathcal L=-\frac1T\sum_t\log p(x_t\mid x_{<t})$ and token perplexity is $e^{\mathcal L}$. Two tokenizers can encode the same text with different $T$, so comparing their per-token perplexities as if the denominator were identical is misleading.

To compare likelihood on the same text, report the evaluation representation and a common normalization such as bits per byte when appropriate: total negative log likelihood divided by $\log 2$ and the number of original bytes. Account for special tokens and document boundaries consistently. This does not remove dataset or preprocessing differences, but it makes the unit explicit.

Whitespace and special tokens are part of the format

A leading space can change segmentation. Unicode normalization can change bytes. A chat template can insert role separators and end-of-turn markers. These are not cosmetic if the model was trained on their token IDs. The same visible sentence embedded in a different template can produce a different conditional distribution.

Special token IDs should be inserted through the intended tokenizer/template interface. Do not assume that writing a string resembling a special marker necessarily produces that special ID. Conversely, an interface that recognizes special strings must distinguish user text from control tokens according to its rules.

Small tests with high diagnostic value

Test encode/decode round trips under the tokenizer's stated normalization policy. Inspect leading/trailing spaces, repeated newlines, emoji, accented text, code indentation, and empty strings. Verify BOS/EOS insertion and padding sides. Record the tokenizer revision with the checkpoint: compatible model weights and incompatible token IDs can produce failures that look like a broken network.

Try it: Can a token ID be interpreted as a point on a numerical number line, so that ID 101 is more similar to 102 than 900?

No. IDs are categorical lookup indices. Similarity, if learned, appears in vectors and model behavior, not in the numerical distance between IDs.

From IDs to an objective

Continue with embedding lookup, then training and evaluation to track exactly which positions contribute to the loss.