Attention Residuals: Let a Layer Choose Its Sources Across Depth
What an ordinary residual stream accumulates
Write a simplified residual stack as $h_{\ell+1}=h_\ell+f_\ell(h_\ell)$. Expanding gives $h_L=h_0+\sum_{\ell<L}f_\ell(h_\ell)$. The learned sublayers can control what they write, but the explicit accumulation coefficients are all one. A later layer receives the accumulated sum rather than independently choosing a coefficient for each earlier write.
The individual writes depend on intermediate states; this expansion does not mean that the sublayers can run independently. It only exposes the additive structure. Normalization and gating inside a real block further affect the effective scale of each contribution.
Move the selection axis
For one token position, collect candidate vectors $v_0,\ldots,v_{m-1}\in\mathbb R^d$ from earlier depth sources, including the embedding source. A learned query $w_\ell\in\mathbb R^d$ at the receiving layer scores normalized keys:
This simplified full-AttnRes description follows Attention Residuals; the official implementation gives the exact block bookkeeping. The query is a learned layer parameter, while the source vectors depend on the input token. Consequently, the weights can vary across tokens even though the query is not computed from a separate current-token vector.
Ordinary token attention compares positions within a sequence. This operation compares depth sources at one position. A causal model still needs causal token mixing. AttnRes does not grant access to future tokens on its own, and it does not make a causal mask unnecessary.
A depth mixture you can inspect
Suppose three source vectors are $v_0=(1,0)$, $v_1=(0,2)$, and $v_2=(1,1)$, and their already-computed scores are $(0,\log2,0)$. Their weights are $(1/4,1/2,1/4)$, giving $h=(0.5,1.25)$. The ordinary unweighted sum of these vectors would be $(2,3)$.
This is a comparison of aggregation rules on fixed candidates, not a claim that trained AttnRes and residual models produce those same candidates. Because inputs to later functions change, replacing the residual mechanism changes the entire learned computation.
Why normalize keys but keep values
A large candidate norm could otherwise dominate a dot-product score simply through scale. RMS normalization makes score selection less directly dependent on magnitude. Keeping unnormalized values allows their magnitude to affect the resulting representation. Thus key normalization and value aggregation serve different purposes.
Even with normalized keys, learned queries can create sharp or diffuse distributions. Softmax guarantees weights sum to one, but the output norm is not fixed: candidates can have different norms and directions. A convex combination can reinforce aligned vectors or cancel opposing ones.
Reduce the number of persistent depth sources
Keeping every earlier sublayer output as a separate candidate costs memory proportional to the number of sources. Scoring all prior sources at every layer creates a quadratic depth term. Block AttnRes groups writes into block summaries and retains a partial summary for the current block. Later layers select among the embedding, completed block summaries, and the available current-block contribution.
The detailed partial-block update matters. Simply averaging a completed model's hidden states once at the end does not implement this architecture: AttnRes changes what each subsequent sublayer consumes. Nor is block aggregation an exact algebraic rearrangement of full AttnRes, because separate values have been combined before the receiver can weight them independently.
As a bookkeeping illustration, 24 writes grouped into blocks of 6 create 4 completed block summaries. The receiver has fewer distinct candidates than if all 24 writes remain available. The reduction trades independent source selection for less storage and selection work. Actual implementations choose their counting unit carefully: “layer,” “sublayer,” and “block” need not refer to the same thing.
What changes in backpropagation?
For fixed candidate vectors and weights, the direct value path gives each source a factor $\alpha_i$ in the incoming gradient. There is also a score path because changing a source changes its normalized key, which changes every softmax weight. Ignoring this second path would train a different operation.
Small weights need not mean a source has no effect on the network: it may influence later candidate vectors through intervening computation. Similarly, a high depth weight is not a complete causal explanation of the model's prediction.
Try it: With a learned query fixed for one layer, must every token receive the same depth weights?
No. The candidate keys depend on each token’s representations. Their dot products with the fixed learned query therefore differ across tokens.
Read a real configuration
The Kimi K3 case study combines block AttnRes with recurrent/attention layers and MoE. These mechanisms select across different axes: earlier depth sources, earlier token representations, and expert functions.