DeltaNet: Write the Error, Not Another Copy
Keep the state orientation explicit
Use column keys $k_t\in\mathbb R^{d_k}$, values $v_t\in\mathbb R^{d_v}$, and $S_t\in\mathbb R^{d_k\times d_v}$. The read is $S_t^\top q_t$. We use this orientation in all recurrent chapters; some papers transpose the state, which moves transition matrices to the other side.
Start from additive writes $S_t=S_{t-1}+k_tv_t^\top$. If the same unit key is written twice with scalar value 3, reading that key returns 6 after the second write. For a “latest value for this key” task, accumulation is the wrong behavior.
Fit the current association with one gradient step
At token $t$, define an auxiliary squared error for the temporary state:
A step with rate $\beta_t$ gives
The error is a value-space vector. Its outer product with the key directs the write to the key's feature direction. This is an interpretation of a forward-pass state update, not an optimizer step on the model's persistent trained weights during inference.
What happens when we read the same key?
Let $\widehat v=S_{t-1}^\top k_t$ and assume $\|k_t\|=1$. Multiply the updated state by that key:
At $\beta_t=1$, the association becomes exactly $v_t$ in this idealized update. At $\beta_t=0.25$, the retrieved value moves one quarter of the way from the old value to the new one. Without unit normalization, the effective step is multiplied by $\|k_t\|^2$; normalization is therefore part of the stability story.
For an old value 3 and a new value 5, a rate of 0.5 yields 4. Writing the same association again yields 4.5, then 4.75. Additive memory would instead grow by five each time. This simple overwrite task is a useful diagnostic because its desired behavior is unambiguous.
A write can affect other keys
For any query $q$, the output change is
A query orthogonal to the written key sees no change. A similar query sees some of the update. If two desired associations use nearly identical keys but require different values, editing one can damage the other. The delta rule improves update behavior; it does not give a finite matrix unlimited independently addressable slots.
For unit keys and $0\le\beta\le1$, $I-\beta kk^\top$ preserves directions orthogonal to $k$ and scales the $k$ direction by $1-\beta$. This explains the selective erasure in the transition form. More general parameterizations require their own stability conditions.
A minimal update
def delta_step(state, key, value, beta, query):
prediction = state.T @ key
error = value - prediction
new_state = state + beta * np.outer(key, error)
output = new_state.T @ query
return new_state, output
The query reads the state after the current token's write in this inclusive causal convention. Reading before writing defines a different alignment. The model's token projections, output gate, normalization, and short convolution are additional components, not hidden inside this equation.
How a sequential recurrence can be trained in chunks
Write each update as $S_t=A_tS_{t-1}+B_t$. Two consecutive updates compose to $(A_2A_1)S_0+(A_2B_1+B_2)$. The composition rule is associative, which provides a route to parallel scans. Naïvely multiplying dense transition matrices, however, can cost too much.
The useful structure is that $A_t=I-\beta_tk_tk_t^\top$ is an identity plus a rank-one update. Parallelizing Linear Transformers with the Delta Rule over Sequence Length develops an efficient chunkwise treatment using this structure. The paper's contribution should not be reduced to “just run the loop in parallel”: data dependencies still have to be represented and computed.
What this chapter's equation leaves open
Constant state size describes decoding at fixed head dimensions. It does not establish a universal throughput advantage, nor does it establish equivalence to softmax attention. The delta-rule read is generally unnormalized and can produce signed outputs; the normalization-vector formula from the preceding chapter is not automatically present.
Try it: With a unit key and beta equal to one, which information is guaranteed to survive the overwrite?
The state components in directions orthogonal to the key survive unchanged in this idealized update. The component along the key is replaced to fit the new association. Overlapping keys can still interfere.
Add controlled forgetting
Gated DeltaNet introduces a separate mechanism for discarding old state, including information orthogonal to the current key.