← All Posts
Deep Learning · Transformers· Position and architecture

RoPE: Relative Position Through Rotated Queries and Keys

RoPE makes the query/key dot product depend on relative phase. It rotates coordinate pairs by position-dependent angles before attention. Values are ordinarily left unrotated, so the mechanism changes matching rather than directly rotating the retrieved content.

Begin with one two-dimensional plane

For a column vector, define the rotation $R(\theta)=\begin{pmatrix}\cos\theta&-\sin\theta\\\sin\theta&\cos\theta\end{pmatrix}$. At position $i$, rotate a query by $R(i\omega)$ and a key at position $j$ by $R(j\omega)$, using the same frequency $\omega$ for that pair.

$$(R(i\omega)q)^\top(R(j\omega)k)=q^\top R((j-i)\omega)k.$$

The identity follows from $R(\theta)^\top=R(-\theta)$ and composition of rotations. The positional part depends on $j-i$. The content vectors $q,k$ still depend on their tokens and context; RoPE does not make an entire attention score a function only of distance.

A numerical relative-phase example

Take $q=k=(1,0)$ and frequency $\omega=\pi/4$. Query position 1 rotates to $(\sqrt2/2,\sqrt2/2)$. Key position 3 rotates to $(-\sqrt2/2,\sqrt2/2)$. Their dot product is zero, equal to $\cos((3-1)\pi/4)$.

Shift both positions forward by five: their difference stays two, so the dot product remains zero. Change only the key position to 2 and the dot product becomes $\cos(\pi/4)=\sqrt2/2$. These facts are exact for the chosen vectors, independent of any model training.

Rotate a query and a key
Change the relative offset. The diagram and dot product use one two-dimensional plane; a real head combines many frequencies.

Apply different clocks to different coordinate pairs

For even rotary dimension $d_r$, use frequencies $\omega_r=\theta^{-2r/d_r}$ for $r=0,\ldots,d_r/2-1$. The rotation matrix is block diagonal, with one 2D rotation per frequency. Some models rotate only part of each head; the remaining coordinates have no rotary transformation.

Interleaved-pair and split-half tensor layouts encode the same idea with different coordinate ordering. Their cached sine/cosine tables and rotation functions must agree. Mixing conventions can produce plausible shapes and completely wrong values.

def rotate_pairs(x, position, frequencies):
    # x: [..., 2*r], with adjacent coordinates forming pairs
    pairs = x.reshape(*x.shape[:-1], -1, 2)
    angle = position * frequencies
    a, b = pairs[..., 0], pairs[..., 1]
    out = np.stack([a*np.cos(angle) - b*np.sin(angle),
                    a*np.sin(angle) + b*np.cos(angle)], axis=-1)
    return out.reshape(x.shape)

Cached keys retain their original positions

During prefill, key $j$ is rotated with its own logical position and can be cached in that form. A new query at position $i$ is rotated with position $i$ and compared directly with those cached keys. Do not rotate already rotated keys again every step.

If the entire sequence is shifted by a common offset and both Q and K rotations are adjusted consistently, the relative dot-product identity is preserved. Changing only the positions of newly generated tokens or resetting them to zero breaks the intended relation.

Why a formula at any position is not unlimited context

Rotations are periodic in each plane. New lengths expose the network to unfamiliar combinations of phases and distances, and attention must distinguish useful matches from competing tokens over a larger candidate set. A model can therefore fail far beyond its training lengths even though every trigonometric value is finite.

Position interpolation rescales positions into a shorter phase range. Changing the base or rescaling selected frequencies changes how different distance scales are represented. Such changes trade local resolution against longer-range behavior and usually need validation or adaptation. The existing RoPE and YaRN chapter develops extension schedules in more detail.

What rotation preserves

Each rotation is orthogonal, so it preserves the norm of the vector being rotated. It does not preserve a dot product between vectors rotated by different angles—that relative change is the point. Shared orthogonal rotation preserves a dot product; different positional rotations introduce the offset-dependent factor above.

This distinction is also central to MLA: position-dependent rotation can obstruct moving a key projection into the query side. Separating positional and content paths allows a compact cache representation.

Try it: If both query and key positions increase by the same amount, what happens to their rotary dot product?

It is unchanged for fixed unrotated content vectors and the same frequency schedule, because only their position difference appears in the identity.

Read the primary formulation

The reference is RoFormer. Compare with ALiBi, which injects a distance preference directly into the logits instead of rotating coordinates.