← All Posts
Deep Learning · Transformers· Foundations

Attention from First Principles: One Complete Calculation

Attention lets one token read information from other tokens. It gives each available position a weight, then uses those weights to combine the information stored there. The result is a vector that later parts of the model use to make a prediction.

Inside a block → attention → one head. View the model diagram →

Give a token some context

Take the prefix “the cat sat”. The embedding for “sat” starts from a table lookup; that lookup alone says nothing about which animal sat. Attention lets the vector at “sat” read information from the other positions. We will follow just one such read.

Each input vector is transformed into three learned vectors with different jobs:

The query is compared with each available key. A larger match score gives the corresponding value more weight. Keys decide how much to read; values decide what is read. These roles are learned through training.

Follow one query through three steps

Use a query $q=(1,0)$ and the small key/value vectors below. These numbers are chosen for the example, not taken from a trained model. Keep the vectors fixed as you follow the scores into weights and then into an output.

One query reads three values
Fixed inputs · query = (1, 0)
PositionKeyValue
1(1, 0)(2, 0)
2(0, 1)(0, 2)
3(1, 1)(2, 2)
  1. Compare the query with each key

    The dot products are (1, 0, 1). Divide by √2 because each key has two features.

    Scores = (0.7071, 0.0000, 0.7071)

  2. Turn scores into weights with softmax

    A bigger score gets a bigger share. The three shares add up to one.

    Position 10.4011
    Position 20.1978
    Position 30.4011
  3. Multiply each value by its weight, then add

    The values are (2, 0), (0, 2), and (2, 2). The bars show how much of each enters the result.

    Output = (1.6044, 1.1978) All three positions contribute.
One read produces one output vector. Hiding a position changes the weights and output while leaving the query, keys, and values fixed.

With all three positions allowed, the final step is:

$$o\approx0.4011(2,0)+0.1978(0,2)+0.4011(2,2)\approx(1.6044,1.1978).$$

Positions 1 and 3 have equal scores but different values. Changing value 3 would change the output without changing the weights. Changing key 3 can change every weight, because the scores compete for a total share of one.

How the scores become shares

Softmax exponentiates the scores, then divides each exponential by their sum. For our scores, subtracting the largest score first gives exponentials $(1,0.4931,1)$. Dividing by their sum gives the weights $(0.4011,0.1978,0.4011)$.

$$a_j=\frac{e^{s_j-c}}{\sum_r e^{s_r-c}},\qquad c=\max_r s_r.$$

Subtracting the same number from every score leaves these ratios unchanged and avoids very large exponentials. Softmax runs across the available keys for one query. Before dropout, the output is a weighted average of their values.

Optional: gradients and the entropy interpretation
$$\frac{\partial a_j}{\partial s_k}=a_j(\mathbf1[j=k]-a_k).$$

Raising one score raises its own weight and lowers the others. When one weight is already near one, many of these derivatives are small, so small score changes have little effect.

Softmax also maximizes $\sum_j a_j s_j+H(a)$ over probability vectors, where $H(a)=-\sum_j a_j\log a_j$. The score term rewards better matches; the entropy term rewards spreading weight. At positive temperature $T$, the corresponding objective uses $T H(a)$. This describes the attention weights, not the model’s final vocabulary probabilities.

Why divide by the square root?

A dot product adds one term per feature. As the number of features grows, its scores can become more spread out, making softmax concentrate heavily on the largest score. Dividing by $\sqrt{d_k}$ compensates for that growth; $d_k$ is the number of features in a query or key.

For independent, zero-mean, unit-variance query and key coordinates, the dot product has variance $d_k$. Dividing by $\sqrt{d_k}$ makes that variance one. Learned vectors need not satisfy these assumptions exactly; this calculation explains the scale choice.

Hide forbidden positions before softmax

Try hiding position 3 in the figure. Its score becomes $-\infty$, whose exponential is zero. The other two weights grow to approximately $(0.6698,0.3302)$, and the output becomes $(1.3395,0.6605)$. Simply zeroing its old weight after softmax would leave the remaining weights summing to less than one.

A causal mask applies this rule to future positions. Every valid query must still have at least one allowed key; ordinary softmax is undefined for an entirely masked row.

Do the same calculation for every token

So far we have followed one query. To process all $n$ tokens together, put their $d$-feature input vectors in the rows of a matrix $X$. Three learned projections produce the queries, keys, and values:

$$Q=XW_Q,\qquad K=XW_K,\qquad V=XW_V.$$

$Q$ and $K$ have shape $n\times d_k$; $V$ has shape $n\times d_v$. The value width $d_v$ can differ from the query/key width. Thus $W_Q$ and $W_K$ have shape $d\times d_k$, while $W_V$ has shape $d\times d_v$.

The three steps from the figure now become three matrix operations:

$$S=\frac{QK^\top}{\sqrt{d_k}}+M,\qquad A=\operatorname{softmax}_{\mathrm{row}}(S),\qquad O=AV.$$
ObjectShapeWhat one row holds
$S$: scores$n\times n$One query’s score for every key
$A$: weights$n\times n$That query’s shares, summing to one
$O$: outputs$n\times d_v$The weighted read $O_i=\sum_j A_{ij}V_j$

The mask $M$ adds zero to allowed scores and $-\infty$ to forbidden scores. Each output row comes from the same calculation you just followed, using that row’s own query.

A transparent reference implementation

import numpy as np

def attention(q, k, v, allowed):
    # q: [nq, dk], k: [nk, dk], v: [nk, dv]
    assert allowed.shape == (len(q), len(k))
    assert allowed.any(axis=-1).all()  # no fully masked row
    scores = q @ k.T / np.sqrt(q.shape[-1])
    scores = np.where(allowed, scores, -np.inf)
    probs = np.exp(scores - scores.max(axis=-1, keepdims=True))
    probs /= probs.sum(axis=-1, keepdims=True)
    return probs @ v, probs

This materializes the full score and probability matrices for clarity. Production kernels can compute the same function without retaining those matrices; see FlashAttention. A function that uses Q @ (K.T @ V) has removed softmax and is a different attention mechanism.

What self-attention does and does not mean

Self-attention means queries, keys, and values come from the same sequence. It does not mean each token attends only to itself. Cross-attention takes queries from one sequence and keys/values from another. Causality and head sharing are separate choices.

Without positional information or a position-dependent mask, permuting the input rows permutes the outputs in the same way. Formally, for a permutation matrix $P$, attention on $PX$ produces $PO$. This is permutation equivariance, not invariance: the order of output rows still follows the input order. Position encodings make relative or absolute location available.

Try it: If every value vector equals the same vector $v$, can changing the query change the output?

Without attention dropout, no: the weights sum to one, so their weighted sum is always $v$. The attention pattern can change while the retrieved representation stays identical.

Continue the construction

Read multi-head attention next. The primary formulation is Vaswani et al., Attention Is All You Need; the arithmetic and counterexamples here are constructed for this tutorial.