Position Encodings: Order Must Enter Somewhere
The permutation argument
Let $P$ permute token rows. If $Q=XW_Q$, then queries for $PX$ are $PQ$; similarly K and V are permuted. The score matrix becomes $P(QK^\top)P^\top$. Row-wise softmax respects this simultaneous row/column permutation, so the output is permuted by $P$ too.
This means self-attention without positional structure is permutation equivariant. It does not mean it returns the same sequence after shuffling: output rows are shuffled in correspondence with input rows. A causal mask already introduces order through visibility, but explicit positional information supplies additional ways to distinguish distances and arrangements.
Add a position vector to each token vector
A learned absolute position table gives each supported position its own trainable vector. The model learns how token and position features should coexist. Addition keeps width $d$ fixed; concatenation would increase width unless another projection reduced it. Addition does not make the two signals uniquely recoverable in every possible vector: the training problem learns a useful combined representation.
A table trained for positions 0 through 1023 has no learned entry at 4096. Extending the allocation alone does not train the new rows. Even a formula defined at arbitrary positions does not guarantee that a model trained on short contexts will use much longer contexts accurately.
Construct a sinusoidal encoding
For even width $d$ and channel pair $r$, define $\omega_r=10000^{-2r/d}$ and
Each pair is a clock rotating at a different frequency. Nearby positions change rapidly in high-frequency pairs and slowly in low-frequency pairs. The combination gives a multiscale signal that subsequent projections can use. These positions are measured in token steps, not characters, words, or seconds.
At width four, the frequencies are 1 and 0.01. Position zero is $(0,1,0,1)$. Position one is approximately $(0.8415,0.5403,0.0100,0.99995)$. Position two is approximately $(0.9093,-0.4161,0.0200,0.99980)$. The slow pair distinguishes long distances with gradual phase changes; the fast pair changes substantially over a few tokens.
Why a relative shift is a linear transformation
Angle addition gives
The matrix depends on the offset $\Delta$, not the starting position $i$. This makes offset relationships accessible to linear maps. It does not mean arbitrary learned attention scores become functions only of relative position: projected token content, cross terms, and arbitrary projection matrices still matter.
Three different ways to expose position
| Mechanism | Where position enters | Main question to check |
|---|---|---|
| Learned or sinusoidal absolute | Added to input vectors | How are positions beyond training represented? |
| RoPE | Rotates query/key coordinates | How do phases behave at the requested distances? |
| ALiBi | Adds distance bias to logits | How strongly are distant positions penalized? |
These approaches are not interchangeable checkpoint options. They change the function a model was trained to compute. A pretrained model expecting learned absolute positions does not become a well-trained RoPE model by replacing one line of inference code.
A reference sinusoidal table
def sinusoidal_positions(length, width):
assert width % 2 == 0
positions = np.arange(length)[:, None]
frequencies = 10000.0 ** (-np.arange(0, width, 2) / width)
phase = positions * frequencies[None, :]
table = np.empty((length, width))
table[:, 0::2] = np.sin(phase)
table[:, 1::2] = np.cos(phase)
return table
Generation with a cache must use absolute positions that continue after the prefix; resetting each new token to position zero changes its relation to cached keys. Packed documents need an explicit position-ID convention alongside their attention boundaries. Padding should not accidentally consume or reset meaningful positions unless the model was trained that way.
Try it: Does a sinusoidal encoding allow evaluation at one million positions?
The formula does. Whether the model remains accurate at those positions is an empirical question about its training distribution, phase behavior, attention, and task. A defined encoding and useful length generalization are different claims.
Continue with rotation
The sinusoidal construction appears in the original Transformer. RoPE uses rotations inside the attention dot product, where a relative-offset identity can be stated more directly.