RoPE from First Principles, and YaRN
The requirement
Attention compares a query at position $m$ with a key at position $n$ through an inner product. We want position to enter that comparison, but we want the result to be translation invariant: shifting both tokens by the same amount should not change their score. Formally, we want functions $f_q,f_k$ such that
for some $g$ depending on position only through the difference $n-m$.
Additive absolute embeddings fail this. Adding $p_m$ to the input makes the score contain cross terms in $m$ and $n$ separately, so the model must learn approximate translation invariance instead of being given it.
Solving it in two dimensions
Start with $d=2$, where a vector is a complex number. Suppose we encode position by a rotation of angle $m\theta$:
Rotation matrices are orthogonal, so $R_\phi^\top=R_{-\phi}$ and $R_\phi R_\psi=R_{\phi+\psi}$. The inner product becomes
Lifting to d dimensions
A single rotation angle can only encode position modulo $2\pi/\theta$. So split the head dimension into $d/2$ independent planes and give each its own frequency. With $\Theta=\{\theta_1,\dots,\theta_{d/2}\}$, the position-$m$ operator is block diagonal:
and because a block-diagonal product of rotations is still orthogonal with the same composition rule, the relative property survives intact:
The frequencies use a geometric schedule borrowed from sinusoidal embeddings, with base $b=10000$:
Each plane has a wavelength $\lambda_i=2\pi/\theta_i$, the number of positions needed for a full revolution. The first plane turns once every $2\pi\approx6.3$ tokens; the last turns once every $2\pi\,b^{(d-2)/d}$ tokens.
Implementing it without building a matrix
Materializing a $d\times d$ block-diagonal matrix would be absurd; it is $99\%$ zeros. Because each block acts on two coordinates, the whole operation is elementwise:
where $\text{swap}$ negates one member of each pair and exchanges them. Two caches of shape $(\text{max positions},d)$ holding $\cos$ and $\sin$ are precomputed once, and applying RoPE costs two multiplies and one add per element.
rotate_half function does. The two are mathematically equivalent — a fixed permutation of the head dimension maps one to the other — but weights trained under one convention produce nonsense under the other. When porting checkpoints between frameworks, this is the first thing to check.Why the design works
No parameters
RoPE adds nothing to learn. It is a fixed, invertible, norm-preserving transform, so it cannot destroy information or add optimization difficulty.
Applied to Q and K only
Values carry content, not position, so they are left alone. Position enters the attention weights but not the payload being mixed.
Long-range decay
Summing many planes whose relative phases spread out with distance makes the expected score fall off as $|n-m|$ grows, giving a mild locality prior for free.
The long-context problem
A model pretrained with context $L$ has only ever seen rotation angles $m\theta_i$ for $m
Three families of fixes exist, and they differ in exactly one respect: which frequencies they are willing to distort.
Position interpolation
The simplest fix squeezes the new range back into the old one. With scale factor $s=L'/L$, replace the position itself:
Now position $L'$ maps to angle $L$, squarely inside the trained range, and nothing extrapolates. The cost is that every frequency is slowed by $s$, including the fast ones that encoded fine local distinctions. Adjacent tokens become harder to tell apart, which shows up as degraded short-range performance and usually requires some fine-tuning to recover.
NTK-aware scaling
The opposite intuition: rather than scaling positions, change the base so that low frequencies stretch a lot and high frequencies barely move.
Because $\theta_i=b^{-2(i-1)/d}$, raising the base multiplies each frequency by a factor that depends on $i$: nearly $1$ for $i=1$ and nearly $1/s$ for $i=d/2$. The exponent $d/(d-2)$ is chosen precisely so the slowest plane ends up stretched by the full factor $s$. This preserves local resolution far better than plain interpolation and often works with no fine-tuning at all — but it is a blunt instrument, since a single knob controls the whole spectrum.
YaRN: interpolate by parts
YaRN makes the choice explicit per plane. The natural question to ask of dimension $i$ is how many full rotations did it complete within the training context?
Many rotations ($r$ large)
The plane wrapped around thousands of times during pretraining, so the model has seen every phase it could ever see. Its information is genuinely relative and local. Interpolating it would destroy fine distinctions for no benefit, so leave it alone.
Fewer than one rotation ($r$ small)
The plane never completed a cycle, so its phase acts as an absolute position marker. Extending the context takes it into angles never seen. It must be interpolated.
A ramp interpolates between the two regimes, with $\alpha=1$ and $\beta=32$ for Llama-family models:
The second half of YaRN
Interpolation makes rotated queries and keys more similar across positions, which flattens the attention distribution. YaRN compensates by sharpening the logits with a temperature that grows logarithmically with the scale factor:
In implementations this factor is folded into the cached $\cos$ and $\sin$ tables, so it multiplies both the query and the key. The logit therefore scales by the square:
At $s=8$ that is a factor of about $1.46$ — a modest but measurable sharpening, obtained for free since the tables were precomputed anyway.
Choosing between them
| Method | What it changes | High frequencies | Fine-tuning needed |
|---|---|---|---|
| Extrapolation (do nothing) | nothing | see unseen phases | fails outright |
| Position interpolation | $m\mapsto m/s$ | squashed by $s$ | yes, to recover local acuity |
| NTK-aware | $b\mapsto b\,s^{d/(d-2)}$ | almost unchanged | often none |
| YaRN | per-plane ramp plus logit temperature | untouched by construction | little; best quality per token of fine-tuning |
RoPE is the unique clean answer to "make the score depend only on relative position", implemented as one rotation per frequency plane. Every context-extension method is a policy for which planes may be slowed down, and YaRN's policy — slow the ones that never completed a rotation, leave the rest alone — is simply the most informed one.
Check yourself
- Prove the relative-position property from orthogonality of $R_\phi$. proof
- Compute the longest wavelength for $d=128$, $b=10000$, and explain why it is not $2\pi b$. calculation
- For $L=4096$ and $d=128$, find which dimension pairs satisfy $r(i)<1$. calculation
- Show that $\gamma=0$ yields $\theta_i/s$ and $\gamma=1$ yields $\theta_i$, and say which frequencies get which. reasoning
- Explain why folding the YaRN temperature into $\cos$ and $\sin$ squares its effect on the logits. derivation
- Implement both pairing conventions and verify they give identical attention scores under the matching permutation. implementation