Four approaches to replacing fixed residual accumulation with learned depth-wise aggregation in Transformers. Watch how information flows differently in each architecture.
Standard Residual Connection
Every layer adds its output with weight 1. No selection, no forgetting. The hidden state grows unboundedly with depth.
DenseFormer
Feb 2024 · EPFL · NeurIPS '24
Static scalar weights · Single stream · O(L²) params
After each block, compute a Depth-Weighted Average: a learned linear combination of all previous layer outputs. Weights are static — learned once during training, then fixed. The simplest version of "don't give every layer weight 1."
Multiway Dynamic Dense connections: generate aggregation weights from hidden states via a small MLP. Four independent channels for Q, K, V, and residual streams — each learns which earlier layers matter for its specific role. 2.4× compute saving.
Three independent Generalized Residual Networks transform inputs for Q, K, V. Each applies nonlinear, input-dependent, dimension-dependent weighting across previous layers. 3× faster convergence to same quality.
Attention Residuals
Mar 2026 · Moonshot AI · Kimi Team
Softmax attention over depth · Single stream · 1 vector per layer
Each layer has a learned pseudo-query vector. It attends over all previous layer outputs via softmax, forcing weights to sum to 1. The simplest formulation: directly prevents magnitude growth. 1.25× compute advantage, <2% inference overhead.