Vision Transformers: Turn an Image into a Sequence of Patches
From pixels to a sequence
Take an image of height $H$, width $W$, and $C$ channels. With non-overlapping square patches of side $P$, assuming exact divisibility, there are $N=(H/P)(W/P)$ patches. Flatten each patch into a vector of length $P^2C$ and apply a shared learned projection $W_E\in\mathbb R^{P^2C\times d}$.
For a $224\times224$ RGB image and $P=16$, there are $14\times14=196$ patches, each with $16\times16\times3=768$ scalar pixel values. With hidden width 768, the patch projection matrix is $768\times768$. Add a classification token and the encoder sequence has 197 positions.
The projection is shared across locations, just as a token embedding system uses a common vocabulary table. It can be implemented as a convolution with kernel size and stride equal to the patch size. This implementation equivalence does not mean that the whole ViT is a convolutional network.
The encoder needs spatial information
If the patch embeddings are merely permuted and there is no positional mechanism, self-attention is permutation-equivariant over those patches. A classification token's readout would not know which patch originally came from the top-left. Add spatial position information so the model can distinguish arrangements.
The original ViT paper uses learned position embeddings in its principal architecture. The official implementation provides model and training details. Other vision models can use relative positions, windows, hierarchies, or different tokenization; those are further design choices.
What changes inside the Transformer?
The central operations are familiar: multi-head self-attention, a token-wise MLP, normalization, and residual connections. For image classification, all observed patches are available, so ordinary bidirectional attention is natural. An autoregressive image model would require a generation order and a different masking/objective setup.
The class token begins as a learned vector and exchanges information with patch tokens through layers. A classification head maps its final representation to class logits. Mean pooling over patch outputs is another design option; it changes the readout and should match training.
Halving patch width quadruples tokens
At the same image resolution, reducing $P$ from 16 to 8 changes the patch grid from $14\times14$ to $28\times28$: 196 patches become 784. Ignoring a class token, that is four times as many tokens and sixteen times as many pairwise attention scores per head.
The smaller patches preserve finer spatial granularity, but they also increase token-wise projections, MLP work, and attention memory. Increasing image resolution at fixed patch size has the same token-count consequence. Report resolution and patch size alongside model width when comparing costs.
Halving patch side quadruples token count and multiplies pairwise attention scores by sixteen, excluding any special token. Enable JavaScript to change the example inputs; the complete calculation remains in the article.
Position tables need a policy when the grid changes
A learned position table trained for a $14\times14$ grid does not directly supply a distinct vector for every position in a $24\times24$ grid. A common adaptation reshapes patch positions into a two-dimensional grid and interpolates it to the new grid, handling special tokens separately.
Interpolation provides an initialization or input adaptation, not a guarantee of unchanged quality. Aspect ratio, preprocessing, crop policy, and any subsequent fine-tuning matter. Do not silently interpolate the classification token as though it were another spatial grid cell.
Patch boundaries create a modeling choice
A tiny object crossing a patch boundary is represented through multiple projected patches. Attention can combine them, but the initial non-overlapping projection does not explicitly impose every local continuity assumption found in a convolutional hierarchy. Data and training must teach useful spatial relationships.
For dense prediction such as segmentation, a single class vector is insufficient. Preserve or reconstruct spatially indexed representations and use an appropriate decoder/head. For a multimodal language model, a vision encoder and projector can produce visual representations consumed by a language stack; this is a larger system than a standalone ViT classifier.
Test the image-to-token boundary
Use a synthetic image where each patch has a distinct constant value. Verify patch order, channel order, and the reconstructed grid. Check that batch layout and image normalization match the checkpoint. A wrong RGB order or a transposed patch grid can damage predictions while every tensor still has a valid shape.
Try it: At fixed image resolution, what happens to attention score count when patch side length doubles?
The number of patches falls by a factor of four, so pairwise score count falls by roughly a factor of sixteen, ignoring special tokens. This also makes each patch spatially coarser.
Recognize the shared backbone
Return to the architecture comparison and position representations. The same attention algebra supports different input modalities once their vectors and positions are defined.