The Vertical Attention Family

Four approaches to replacing fixed residual accumulation with learned depth-wise aggregation in Transformers. Watch how information flows differently in each architecture.

Standard Residual Connection

Every layer adds its output with weight 1. No selection, no forgetting. The hidden state grows unboundedly with depth.

DenseFormer

Feb 2024 · EPFL · NeurIPS '24
Static scalar weights · Single stream · O(L²) params
After each block, compute a Depth-Weighted Average: a learned linear combination of all previous layer outputs. Weights are static — learned once during training, then fixed. The simplest version of "don't give every layer weight 1."

MUDDFormer

Feb 2025 · Caiyun AI / BUPT · ICML '25
Dynamic weights · Multiway (Q, K, V, residual) · 4-head depth attention
Multiway Dynamic Dense connections: generate aggregation weights from hidden states via a small MLP. Four independent channels for Q, K, V, and residual streams — each learns which earlier layers matter for its specific role. 2.4× compute saving.

DeepCrossAttention

Feb 2025 · Google Research · ICML '25
Input-dependent, dimension-wise weights · Multiway (Q, K, V) · GRN-based
Three independent Generalized Residual Networks transform inputs for Q, K, V. Each applies nonlinear, input-dependent, dimension-dependent weighting across previous layers. 3× faster convergence to same quality.

Attention Residuals

Mar 2026 · Moonshot AI · Kimi Team
Softmax attention over depth · Single stream · 1 vector per layer
Each layer has a learned pseudo-query vector. It attends over all previous layer outputs via softmax, forcing weights to sum to 1. The simplest formulation: directly prevents magnitude growth. 1.25× compute advantage, <2% inference overhead.