← All Posts
Deep Learning · Transformers· Position and architecture

ALiBi: A Distance Bias in Attention Logits

ALiBi adds a head-specific preference for nearby keys. Content still determines the query/key score. Distance contributes an additive penalty before softmax, with different heads using different penalty slopes.

A causal distance penalty

For query position $i$ and allowed key $j\le i$, let head $h$ use positive slope $m_h$. Its score is

$$s_{ij}^{(h)}=\frac{q_i^{(h)\top}k_j^{(h)}}{\sqrt{d_h}}-m_h(i-j).$$

The causal mask still excludes $j>i$; the distance bias does not replace that mask. The original design uses a specified schedule of slopes across heads, so some heads are more strongly local than others. Exact reproduction should use the paper or implementation's slope schedule, especially when the head count is not a power of two.

How much does a bias change the odds?

Take two keys with equal content scores, one at distance one and one at distance five. With $m=0.5$, their biased scores differ by $0.5(5-1)=2$ in favor of the nearer key. Their softmax probability ratio is $e^2\approx7.39$.

If those are the only keys, the weights are approximately $(0.8808,0.1192)$. If other keys are present, the individual probabilities change because the normalizing denominator changes, but this pairwise ratio stays $e^2$. A distant key can overcome the penalty with a sufficiently larger content score; ALiBi does not impose a hard window.

A multiplicative prior inside the unnormalized weight

$$\exp(s_{ij})=\exp(\text{content score}_{ij})\exp[-m(i-j)].$$

The linear logit penalty becomes exponential decay in the unnormalized attention weight. This gives an interpretable distance scale: an extra $\log2/m$ positions halves the unnormalized weight when content scores are held fixed. For $m=0.5$, that distance is about 1.386 tokens.

This interpretation applies within one attention row. It is not a claim that the model's end-to-end influence from a token decays at that exact rate: heads, residual paths, and layers can combine information in more complex ways.

Use the actual position difference

distance = query_positions[:, None] - key_positions[None, :]
allowed = distance >= 0
bias = -slope * distance
scores = content_scores + bias
scores = np.where(allowed, scores, -np.inf)

This handles cached queries with absolute position offsets. For a bidirectional model, one might use absolute distance or directional variants, but that is a different construction from the causal formula above. State the variant rather than copying a causal implementation into an encoder.

ALiBi and RoPE encode different relationships

ALiBi adds a monotone distance penalty independent of token content. RoPE changes the content dot product through relative rotations and multiple frequencies. Learned absolute embeddings add a position vector at the input. These mechanisms can encourage different extrapolation behavior, but none proves accurate retrieval at arbitrary context lengths.

Train Short, Test Long reports length-extrapolation experiments for ALiBi under its training setup. Treat those results as evidence for the tested configurations, not as a universal guarantee that any existing transformer improves when its position mechanism is swapped.

What memory it does not remove

A distance bias can be generated inside a tiled kernel without storing a full bias matrix. It does not by itself remove token pairs from attention or compress historical keys and values. Dense ALiBi attention still has dense pairwise arithmetic and a growing autoregressive K/V cache.

Try it: Does a distance penalty of minus infinity at every position except the current one describe ordinary ALiBi?

No. That would impose a hard attention restriction. ALiBi uses finite linear penalties on allowed positions, so distant content can still receive weight.

Connect position to efficiency

Read KV caching for logical position handling during generation, or sparse attention for actual removal of distant interactions.