← All Posts
Deep Learning · Transformers· Position and architecture

Position Encodings: Order Must Enter Somewhere

Token identity and token position answer different questions. Content-only self-attention can compare words, but it has no intrinsic notion of which unmasked word came first. Position enters through additional input features, attention-score structure, or a sequence-dependent operation.

The permutation argument

Let $P$ permute token rows. If $Q=XW_Q$, then queries for $PX$ are $PQ$; similarly K and V are permuted. The score matrix becomes $P(QK^\top)P^\top$. Row-wise softmax respects this simultaneous row/column permutation, so the output is permuted by $P$ too.

This means self-attention without positional structure is permutation equivariant. It does not mean it returns the same sequence after shuffling: output rows are shuffled in correspondence with input rows. A causal mask already introduces order through visibility, but explicit positional information supplies additional ways to distinguish distances and arrangements.

Add a position vector to each token vector

$$x_i=E[\mathrm{token}_i]+P[i].$$

A learned absolute position table gives each supported position its own trainable vector. The model learns how token and position features should coexist. Addition keeps width $d$ fixed; concatenation would increase width unless another projection reduced it. Addition does not make the two signals uniquely recoverable in every possible vector: the training problem learns a useful combined representation.

A table trained for positions 0 through 1023 has no learned entry at 4096. Extending the allocation alone does not train the new rows. Even a formula defined at arbitrary positions does not guarantee that a model trained on short contexts will use much longer contexts accurately.

Construct a sinusoidal encoding

For even width $d$ and channel pair $r$, define $\omega_r=10000^{-2r/d}$ and

$$P(i)_{2r}=\sin(i\omega_r),\qquad P(i)_{2r+1}=\cos(i\omega_r).$$

Each pair is a clock rotating at a different frequency. Nearby positions change rapidly in high-frequency pairs and slowly in low-frequency pairs. The combination gives a multiscale signal that subsequent projections can use. These positions are measured in token steps, not characters, words, or seconds.

At width four, the frequencies are 1 and 0.01. Position zero is $(0,1,0,1)$. Position one is approximately $(0.8415,0.5403,0.0100,0.99995)$. Position two is approximately $(0.9093,-0.4161,0.0200,0.99980)$. The slow pair distinguishes long distances with gradual phase changes; the fast pair changes substantially over a few tokens.

Why a relative shift is a linear transformation

Angle addition gives

$$\begin{pmatrix}\sin((i+\Delta)\omega)\\\cos((i+\Delta)\omega)\end{pmatrix}=\begin{pmatrix}\cos(\Delta\omega)&\sin(\Delta\omega)\\-\sin(\Delta\omega)&\cos(\Delta\omega)\end{pmatrix}\begin{pmatrix}\sin(i\omega)\\\cos(i\omega)\end{pmatrix}.$$

The matrix depends on the offset $\Delta$, not the starting position $i$. This makes offset relationships accessible to linear maps. It does not mean arbitrary learned attention scores become functions only of relative position: projected token content, cross terms, and arbitrary projection matrices still matter.

Three different ways to expose position

MechanismWhere position entersMain question to check
Learned or sinusoidal absoluteAdded to input vectorsHow are positions beyond training represented?
RoPERotates query/key coordinatesHow do phases behave at the requested distances?
ALiBiAdds distance bias to logitsHow strongly are distant positions penalized?

These approaches are not interchangeable checkpoint options. They change the function a model was trained to compute. A pretrained model expecting learned absolute positions does not become a well-trained RoPE model by replacing one line of inference code.

A reference sinusoidal table

def sinusoidal_positions(length, width):
    assert width % 2 == 0
    positions = np.arange(length)[:, None]
    frequencies = 10000.0 ** (-np.arange(0, width, 2) / width)
    phase = positions * frequencies[None, :]
    table = np.empty((length, width))
    table[:, 0::2] = np.sin(phase)
    table[:, 1::2] = np.cos(phase)
    return table

Generation with a cache must use absolute positions that continue after the prefix; resetting each new token to position zero changes its relation to cached keys. Packed documents need an explicit position-ID convention alongside their attention boundaries. Padding should not accidentally consume or reset meaningful positions unless the model was trained that way.

Try it: Does a sinusoidal encoding allow evaluation at one million positions?

The formula does. Whether the model remains accurate at those positions is an empirical question about its training distribution, phase behavior, attention, and task. A defined encoding and useful length generalization are different claims.

Continue with rotation

The sinusoidal construction appears in the original Transformer. RoPE uses rotations inside the attention dot product, where a relative-offset identity can be stated more directly.