← All Posts
Deep Learning · Transformers· Recurrent and hybrid models

Gated DeltaNet: Forget Globally, Update Selectively

The delta rate and the forget gate have different jobs. The forget gate scales the existing memory. The delta rate controls how strongly the current association is fitted after that decay. Using one name for both hides an important part of the model.

Why selective overwriting is not enough

DeltaNet can revise the association along the current key direction while preserving orthogonal directions. But sometimes an entire old context should lose relevance: a new document, a changed topic, or a reset in a structured task. A write along one key cannot instantly erase arbitrary information in all other directions.

Introduce a scalar gate $\alpha_t\in(0,1)$, typically computed from the current input, and retain the rate $\beta_t$. With our key-by-value state orientation, the gated update is

$$S_t=(I-\beta_tk_tk_t^\top)(\alpha_tS_{t-1})+\beta_tk_tv_t^\top.$$

This is the transpose-equivalent of the formulation in Gated Delta Networks. We use the same orientation as the preceding chapter so matrix sides remain consistent.

Decay first, then compute the write error

Define $\overline S_t=\alpha_tS_{t-1}$. Read the decayed prediction $\widehat v_t=\overline S_t^\top k_t$, then write $\beta_tk_t(v_t-\widehat v_t)^\top$. The subtraction must use the decayed state.

def gated_delta_step(state, key, value, alpha, beta):
    decayed = alpha * state
    error = value - decayed.T @ key
    return decayed + beta * np.outer(key, error)

The superficially similar expression $\alpha S+\beta k(v-S^\top k)^\top$ is generally different: its error uses the old state. For a scalar old state 4, new value 10, $\alpha=0.5$, and $\beta=0.25$, the intended update gives $2+0.25(10-2)=4$. The other expression gives $2+0.25(10-4)=3.5$.

Read the limiting cases

Gate settingBehavior
$\alpha=1$Ordinary delta-rule update
$\alpha\to0$Old state is erased; the new rank-one write remains
$\beta=0$Decay without learning the current association
$\alpha=1,\beta=0$Hold the state unchanged

Some implementations parameterize gates with sigmoid-like functions, so exact zero and one may be limiting cases rather than attained values. The table explains their mathematical roles, not the precise activation used by every checkpoint.

A retention curve you can calculate

With no writes and constant $\alpha$, $S_t=\alpha^tS_0$. At $\alpha=0.9$, after ten steps only $0.9^{10}\approx0.349$ of the initial state remains. The half-life is $\log(0.5)/\log(\alpha)$, approximately 6.58 steps for this gate.

With input-dependent gates, retention along an interval is a product $\prod_t\alpha_t$. A single very small gate can dominate that product. This gives flexible reset behavior but also means that a long context window does not imply every early detail remains available in the recurrent state.

Explore retention and new writes
Change the scalar retention gate and the delta rate independently. The example keeps the state and the requested association visible at each step.

Scalar decay cannot choose a feature direction

Multiplying by $\alpha_t$ scales every key-feature row of the state equally. That is useful for a broad reset but coarse when some information should be retained and other information forgotten. The delta correction is direction-specific, yet the decay itself remains uniform within the head.

This distinction motivates Kimi Delta Attention, which replaces the scalar decay with a vector of retention factors. A per-channel gate is a meaningful generalization because the decay and overwrite transformations then need not commute.

What makes the architecture more than a recurrence

A practical token mixer also constructs queries, keys, values, and gates from the residual stream, may apply local convolutions and normalization, and maps the recurrent read back into the model width. Training uses chunkwise algorithms so long sequences do not require a literal Python loop for every token and every layer.

The affine-composition view from DeltaNet still applies, but now each transition includes decay. Hardware efficiency depends on exploiting that structure, choosing chunk sizes, and controlling intermediate numerical ranges. A fixed-size state and a well-designed parallel training kernel address different phases of execution.

Try it: If beta is zero but alpha is 0.8, does the memory remain unchanged?

No. There is no new association write, but the old state is multiplied by 0.8 at each step. The write rate and the retention gate are independent controls.

Give the forget gate more resolution

Continue to Kimi Delta Attention, then hybrid architectures to see how recurrent layers coexist with token-addressable attention.