← All Posts
Deep Learning · Popular Videos · Umar Jamil· Part 2 · Chapters 3–4

RoPE from First Principles, and YaRN

One requirement generates the whole design. Ask that the attention score between two tokens depend only on how far apart they are, and rotary embeddings fall out as essentially the only answer. Every long-context trick that follows is a different way of stretching the frequencies that requirement produced.

The requirement

Attention compares a query at position $m$ with a key at position $n$ through an inner product. We want position to enter that comparison, but we want the result to be translation invariant: shifting both tokens by the same amount should not change their score. Formally, we want functions $f_q,f_k$ such that

$$\langle f_q(\mathbf x_m,m),\,f_k(\mathbf x_n,n)\rangle=g(\mathbf x_m,\mathbf x_n,\,n-m)$$

for some $g$ depending on position only through the difference $n-m$.

Additive absolute embeddings fail this. Adding $p_m$ to the input makes the score contain cross terms in $m$ and $n$ separately, so the model must learn approximate translation invariance instead of being given it.

Solving it in two dimensions

Start with $d=2$, where a vector is a complex number. Suppose we encode position by a rotation of angle $m\theta$:

$$f_q(\mathbf x_m,m)=R_{m\theta}\,\mathbf q_m,\qquad R_\phi=\begin{pmatrix}\cos\phi&-\sin\phi\\ \sin\phi&\cos\phi\end{pmatrix}.$$

Rotation matrices are orthogonal, so $R_\phi^\top=R_{-\phi}$ and $R_\phi R_\psi=R_{\phi+\psi}$. The inner product becomes

$$\big(R_{m\theta}\mathbf q\big)^\top\big(R_{n\theta}\mathbf k\big) =\mathbf q^\top R_{m\theta}^\top R_{n\theta}\,\mathbf k =\mathbf q^\top R_{(n-m)\theta}\,\mathbf k.$$
This is the entire idea. Absolute rotations applied separately to the query and the key produce a score that depends only on the relative rotation $(n-m)\theta$. The group structure of rotations does all the work: composing them adds angles, and transposing one negates its angle.

Lifting to d dimensions

A single rotation angle can only encode position modulo $2\pi/\theta$. So split the head dimension into $d/2$ independent planes and give each its own frequency. With $\Theta=\{\theta_1,\dots,\theta_{d/2}\}$, the position-$m$ operator is block diagonal:

$$\mathbf R^d_{\Theta,m}=\begin{pmatrix}R_{m\theta_1}&&&\\&R_{m\theta_2}&&\\&&\ddots&\\&&&R_{m\theta_{d/2}}\end{pmatrix},$$

and because a block-diagonal product of rotations is still orthogonal with the same composition rule, the relative property survives intact:

$$\boxed{\big(\mathbf R^d_{\Theta,m}\mathbf q\big)^\top\big(\mathbf R^d_{\Theta,n}\mathbf k\big)=\mathbf q^\top\mathbf R^d_{\Theta,\,n-m}\,\mathbf k.}$$

The frequencies use a geometric schedule borrowed from sinusoidal embeddings, with base $b=10000$:

$$\theta_i=b^{-2(i-1)/d},\qquad i=1,\dots,d/2.$$

Each plane has a wavelength $\lambda_i=2\pi/\theta_i$, the number of positions needed for a full revolution. The first plane turns once every $2\pi\approx6.3$ tokens; the last turns once every $2\pi\,b^{(d-2)/d}$ tokens.

A detail that trips people up. The slowest frequency is $b^{-(d-2)/d}$, not $b^{-1}$, because the exponent runs to $2(d/2-1)/d$ rather than to $1$. For $d=128$ that gives a longest wavelength of about 54,000 tokens rather than 63,000. It matters when you compute which dimensions YaRN will interpolate.

Implementing it without building a matrix

Materializing a $d\times d$ block-diagonal matrix would be absurd; it is $99\%$ zeros. Because each block acts on two coordinates, the whole operation is elementwise:

$$\mathbf R^d_{\Theta,m}\mathbf x=\mathbf x\odot\cos(m\boldsymbol\theta)+\text{swap}(\mathbf x)\odot\sin(m\boldsymbol\theta),$$

where $\text{swap}$ negates one member of each pair and exchanges them. Two caches of shape $(\text{max positions},d)$ holding $\cos$ and $\sin$ are precomputed once, and applying RoPE costs two multiplies and one add per element.

Two incompatible pairing conventions exist. The original paper pairs adjacent coordinates: $(x_1,x_2),(x_3,x_4),\dots$ The implementation used by Llama and most of the HuggingFace ecosystem instead pairs coordinate $i$ with coordinate $i+d/2$, which is what the widely copied rotate_half function does. The two are mathematically equivalent — a fixed permutation of the head dimension maps one to the other — but weights trained under one convention produce nonsense under the other. When porting checkpoints between frameworks, this is the first thing to check.

Why the design works

No parameters

RoPE adds nothing to learn. It is a fixed, invertible, norm-preserving transform, so it cannot destroy information or add optimization difficulty.

Applied to Q and K only

Values carry content, not position, so they are left alone. Position enters the attention weights but not the payload being mixed.

Long-range decay

Summing many planes whose relative phases spread out with distance makes the expected score fall off as $|n-m|$ grows, giving a mild locality prior for free.

The long-context problem

A model pretrained with context $L$ has only ever seen rotation angles $m\theta_i$ for $mextrapolation failure, and it is severe rather than graceful.

Three families of fixes exist, and they differ in exactly one respect: which frequencies they are willing to distort.

Position interpolation

The simplest fix squeezes the new range back into the old one. With scale factor $s=L'/L$, replace the position itself:

$$m\ \longmapsto\ \frac{m}{s},\qquad\text{equivalently}\qquad \theta_i\ \longmapsto\ \frac{\theta_i}{s}.$$

Now position $L'$ maps to angle $L$, squarely inside the trained range, and nothing extrapolates. The cost is that every frequency is slowed by $s$, including the fast ones that encoded fine local distinctions. Adjacent tokens become harder to tell apart, which shows up as degraded short-range performance and usually requires some fine-tuning to recover.

NTK-aware scaling

The opposite intuition: rather than scaling positions, change the base so that low frequencies stretch a lot and high frequencies barely move.

$$b'=b\cdot s^{\,d/(d-2)}.$$

Because $\theta_i=b^{-2(i-1)/d}$, raising the base multiplies each frequency by a factor that depends on $i$: nearly $1$ for $i=1$ and nearly $1/s$ for $i=d/2$. The exponent $d/(d-2)$ is chosen precisely so the slowest plane ends up stretched by the full factor $s$. This preserves local resolution far better than plain interpolation and often works with no fine-tuning at all — but it is a blunt instrument, since a single knob controls the whole spectrum.

YaRN: interpolate by parts

YaRN makes the choice explicit per plane. The natural question to ask of dimension $i$ is how many full rotations did it complete within the training context?

$$r(i)=\frac{L}{\lambda_i}=\frac{L\,\theta_i}{2\pi}.$$

Many rotations ($r$ large)

The plane wrapped around thousands of times during pretraining, so the model has seen every phase it could ever see. Its information is genuinely relative and local. Interpolating it would destroy fine distinctions for no benefit, so leave it alone.

Fewer than one rotation ($r$ small)

The plane never completed a cycle, so its phase acts as an absolute position marker. Extending the context takes it into angles never seen. It must be interpolated.

A ramp interpolates between the two regimes, with $\alpha=1$ and $\beta=32$ for Llama-family models:

$$\gamma(r)=\begin{cases}0,& r<\alpha\\[2pt] 1,& r>\beta\\[2pt] \dfrac{r-\alpha}{\beta-\alpha},&\text{otherwise,}\end{cases}$$
$$\boxed{\theta_i'=\big(1-\gamma(r(i))\big)\frac{\theta_i}{s}+\gamma(r(i))\,\theta_i.}$$
Get the direction right. $\gamma=0$ means fully interpolated ($\theta_i/s$) and applies to the slow, low-frequency planes. $\gamma=1$ means untouched ($\theta_i$) and applies to the fast, high-frequency planes. It is easy to read the formula backwards, and reversing it produces a model that is worse than plain position interpolation.

The second half of YaRN

Interpolation makes rotated queries and keys more similar across positions, which flattens the attention distribution. YaRN compensates by sharpening the logits with a temperature that grows logarithmically with the scale factor:

$$\sqrt{\tfrac1t}=0.1\ln(s)+1.$$

In implementations this factor is folded into the cached $\cos$ and $\sin$ tables, so it multiplies both the query and the key. The logit therefore scales by the square:

$$\text{logit}\ \longmapsto\ \big(0.1\ln s+1\big)^2\cdot\text{logit}.$$

At $s=8$ that is a factor of about $1.46$ — a modest but measurable sharpening, obtained for free since the tables were precomputed anyway.

original wavelength after scaling training context $L$ target context $L'$
planes untouched
planes in the ramp
planes fully interpolated
logit sharpening
Head dimension 128, base 10000, training context $L=4096$, YaRN thresholds $\alpha=1$, $\beta=32$. The horizontal axis is the dimension-pair index; the vertical axis is wavelength in tokens on a log scale.

Choosing between them

MethodWhat it changesHigh frequenciesFine-tuning needed
Extrapolation (do nothing)nothingsee unseen phasesfails outright
Position interpolation$m\mapsto m/s$squashed by $s$yes, to recover local acuity
NTK-aware$b\mapsto b\,s^{d/(d-2)}$almost unchangedoften none
YaRNper-plane ramp plus logit temperatureuntouched by constructionlittle; best quality per token of fine-tuning
Takeaway

RoPE is the unique clean answer to "make the score depend only on relative position", implemented as one rotation per frequency plane. Every context-extension method is a policy for which planes may be slowed down, and YaRN's policy — slow the ones that never completed a rotation, leave the rest alone — is simply the most informed one.

Check yourself