Gated DeltaNet: Forget Globally, Update Selectively
Why selective overwriting is not enough
DeltaNet can revise the association along the current key direction while preserving orthogonal directions. But sometimes an entire old context should lose relevance: a new document, a changed topic, or a reset in a structured task. A write along one key cannot instantly erase arbitrary information in all other directions.
Introduce a scalar gate $\alpha_t\in(0,1)$, typically computed from the current input, and retain the rate $\beta_t$. With our key-by-value state orientation, the gated update is
This is the transpose-equivalent of the formulation in Gated Delta Networks. We use the same orientation as the preceding chapter so matrix sides remain consistent.
Decay first, then compute the write error
Define $\overline S_t=\alpha_tS_{t-1}$. Read the decayed prediction $\widehat v_t=\overline S_t^\top k_t$, then write $\beta_tk_t(v_t-\widehat v_t)^\top$. The subtraction must use the decayed state.
def gated_delta_step(state, key, value, alpha, beta):
decayed = alpha * state
error = value - decayed.T @ key
return decayed + beta * np.outer(key, error)
The superficially similar expression $\alpha S+\beta k(v-S^\top k)^\top$ is generally different: its error uses the old state. For a scalar old state 4, new value 10, $\alpha=0.5$, and $\beta=0.25$, the intended update gives $2+0.25(10-2)=4$. The other expression gives $2+0.25(10-4)=3.5$.
Read the limiting cases
| Gate setting | Behavior |
|---|---|
| $\alpha=1$ | Ordinary delta-rule update |
| $\alpha\to0$ | Old state is erased; the new rank-one write remains |
| $\beta=0$ | Decay without learning the current association |
| $\alpha=1,\beta=0$ | Hold the state unchanged |
Some implementations parameterize gates with sigmoid-like functions, so exact zero and one may be limiting cases rather than attained values. The table explains their mathematical roles, not the precise activation used by every checkpoint.
A retention curve you can calculate
With no writes and constant $\alpha$, $S_t=\alpha^tS_0$. At $\alpha=0.9$, after ten steps only $0.9^{10}\approx0.349$ of the initial state remains. The half-life is $\log(0.5)/\log(\alpha)$, approximately 6.58 steps for this gate.
With input-dependent gates, retention along an interval is a product $\prod_t\alpha_t$. A single very small gate can dominate that product. This gives flexible reset behavior but also means that a long context window does not imply every early detail remains available in the recurrent state.
Scalar decay cannot choose a feature direction
Multiplying by $\alpha_t$ scales every key-feature row of the state equally. That is useful for a broad reset but coarse when some information should be retained and other information forgotten. The delta correction is direction-specific, yet the decay itself remains uniform within the head.
This distinction motivates Kimi Delta Attention, which replaces the scalar decay with a vector of retention factors. A per-channel gate is a meaningful generalization because the decay and overwrite transformations then need not commute.
What makes the architecture more than a recurrence
A practical token mixer also constructs queries, keys, values, and gates from the residual stream, may apply local convolutions and normalization, and maps the recurrent read back into the model width. Training uses chunkwise algorithms so long sequences do not require a literal Python loop for every token and every layer.
The affine-composition view from DeltaNet still applies, but now each transition includes decay. Hardware efficiency depends on exploiting that structure, choosing chunk sizes, and controlling intermediate numerical ranges. A fixed-size state and a well-designed parallel training kernel address different phases of execution.
Try it: If beta is zero but alpha is 0.8, does the memory remain unchanged?
No. There is no new association write, but the old state is multiplied by 0.8 at each step. The write rate and the retention gate are independent controls.
Give the forget gate more resolution
Continue to Kimi Delta Attention, then hybrid architectures to see how recurrent layers coexist with token-addressable attention.