← All Posts
Deep Learning · Transformers· Recurrent and hybrid models

DeltaNet: Write the Error, Not Another Copy

An additive memory keeps adding even when a key already has a value. The delta rule first reads the old value associated with the current key, then writes only the discrepancy. This turns a memory accumulator into an editable associative state.

Keep the state orientation explicit

Use column keys $k_t\in\mathbb R^{d_k}$, values $v_t\in\mathbb R^{d_v}$, and $S_t\in\mathbb R^{d_k\times d_v}$. The read is $S_t^\top q_t$. We use this orientation in all recurrent chapters; some papers transpose the state, which moves transition matrices to the other side.

Start from additive writes $S_t=S_{t-1}+k_tv_t^\top$. If the same unit key is written twice with scalar value 3, reading that key returns 6 after the second write. For a “latest value for this key” task, accumulation is the wrong behavior.

Fit the current association with one gradient step

At token $t$, define an auxiliary squared error for the temporary state:

$$\mathcal E_t(S)=\frac12\|S^\top k_t-v_t\|^2,\qquad\nabla_S\mathcal E_t=k_t(S^\top k_t-v_t)^\top.$$

A step with rate $\beta_t$ gives

$$S_t=S_{t-1}+\beta_tk_t(v_t-S_{t-1}^\top k_t)^\top=(I-\beta_tk_tk_t^\top)S_{t-1}+\beta_tk_tv_t^\top.$$

The error is a value-space vector. Its outer product with the key directs the write to the key's feature direction. This is an interpretation of a forward-pass state update, not an optimizer step on the model's persistent trained weights during inference.

What happens when we read the same key?

Let $\widehat v=S_{t-1}^\top k_t$ and assume $\|k_t\|=1$. Multiply the updated state by that key:

$$S_t^\top k_t=(1-\beta_t)\widehat v+\beta_tv_t.$$

At $\beta_t=1$, the association becomes exactly $v_t$ in this idealized update. At $\beta_t=0.25$, the retrieved value moves one quarter of the way from the old value to the new one. Without unit normalization, the effective step is multiplied by $\|k_t\|^2$; normalization is therefore part of the stability story.

For an old value 3 and a new value 5, a rate of 0.5 yields 4. Writing the same association again yields 4.5, then 4.75. Additive memory would instead grow by five each time. This simple overwrite task is a useful diagnostic because its desired behavior is unambiguous.

Watch an associative memory update
Compare additive writes with delta-rule writes for repeated scalar associations. The displayed residual is computed before each update.

A write can affect other keys

For any query $q$, the output change is

$$\Delta o=\beta_t(q^\top k_t)(v_t-\widehat v).$$

A query orthogonal to the written key sees no change. A similar query sees some of the update. If two desired associations use nearly identical keys but require different values, editing one can damage the other. The delta rule improves update behavior; it does not give a finite matrix unlimited independently addressable slots.

For unit keys and $0\le\beta\le1$, $I-\beta kk^\top$ preserves directions orthogonal to $k$ and scales the $k$ direction by $1-\beta$. This explains the selective erasure in the transition form. More general parameterizations require their own stability conditions.

A minimal update

def delta_step(state, key, value, beta, query):
    prediction = state.T @ key
    error = value - prediction
    new_state = state + beta * np.outer(key, error)
    output = new_state.T @ query
    return new_state, output

The query reads the state after the current token's write in this inclusive causal convention. Reading before writing defines a different alignment. The model's token projections, output gate, normalization, and short convolution are additional components, not hidden inside this equation.

How a sequential recurrence can be trained in chunks

Write each update as $S_t=A_tS_{t-1}+B_t$. Two consecutive updates compose to $(A_2A_1)S_0+(A_2B_1+B_2)$. The composition rule is associative, which provides a route to parallel scans. Naïvely multiplying dense transition matrices, however, can cost too much.

The useful structure is that $A_t=I-\beta_tk_tk_t^\top$ is an identity plus a rank-one update. Parallelizing Linear Transformers with the Delta Rule over Sequence Length develops an efficient chunkwise treatment using this structure. The paper's contribution should not be reduced to “just run the loop in parallel”: data dependencies still have to be represented and computed.

What this chapter's equation leaves open

Constant state size describes decoding at fixed head dimensions. It does not establish a universal throughput advantage, nor does it establish equivalence to softmax attention. The delta-rule read is generally unnormalized and can produce signed outputs; the normalization-vector formula from the preceding chapter is not automatically present.

Try it: With a unit key and beta equal to one, which information is guaranteed to survive the overwrite?

The state components in directions orthogonal to the key survive unchanged in this idealized update. The component along the key is replaced to fit the new association. Overlapping keys can still interfere.

Add controlled forgetting

Gated DeltaNet introduces a separate mechanism for discarding old state, including information orthogonal to the current key.