Why do transformers need residual connections?
In 2015, a 56-layer convolutional network trained on CIFAR-10 did worse than a 20-layer one. Not on the test set: on the training set. It was not overfitting. It was failing to fit, with more parameters and strictly more capacity, because the deeper network could not even learn to be the shallower one with extra identity layers.
That single plot in the ResNet paper is the reason the residual connection exists, and the transformer inherited it two years later without changing a thing.
The transformer also inherited a normalization layer, and then spent five years arguing about where to put it. I want to answer three interview questions in this part, and they turn out to be the same question asked three ways: why residuals, why LayerNorm rather than BatchNorm, and why the norm moved from after the sublayer to before it. All three are about keeping the gradient alive on its way down a hundred layers.
The 30-second version
With $h_{l+1} = h_l + f_l(h_l)$, the gradient from the output back to layer $l$ is $\prod_{i \ge l}(I + J_i)$, which expands to the identity plus a sum of higher-order terms. The identity term is a direct path from the loss to every layer that no Jacobian can shrink. Without residuals the gradient is $\prod J_i$, a product of $L$ matrices, whose norm grows or decays exponentially in $L$ unless every Jacobian is tuned to have singular values near one. Residuals also make each layer a small perturbation of the identity, so the network at initialization is close to a shallow one and gets deeper as it trains, rather than starting deep and random.What does the skip connection do to the gradient?
A plain stack of layers is $h_{l+1} = f_l(h_l)$. To update layer $l$ you need $\partial \mathcal{L} / \partial h_l$, and the chain rule gives it as a product:
\[\frac{\partial h_L}{\partial h_l} = \prod_{i=l}^{L-1} J_i, \qquad J_i = \frac{\partial f_i}{\partial h_i}\]A product of $L - l$ matrices. If their typical singular value is $s$, the product scales like $s^{L-l}$. For $s = 0.9$ and fifty layers that is $0.005$; for $s = 1.1$ it is $117$. The only stable case is $s = 1$ exactly, in every direction, in every layer, and nothing about a randomly initialized nonlinear layer guarantees that. This is the vanishing and exploding gradient problem, and depth makes it exponentially worse.
Now add the skip. With $h_{l+1} = h_l + f_l(h_l)$:
\[\frac{\partial h_L}{\partial h_l} = \prod_{i=l}^{L-1} (I + J_i) = I + \sum_i J_i + \sum_{i<j} J_j J_i + \cdots\]The first term is the identity. It does not depend on any $J$, it does not shrink with depth, and it carries the gradient from the loss directly to layer $l$ untouched. All the other terms are still there, still products of Jacobians, but they are additions to a baseline of one rather than the whole story. If every $f_i$ were zero, the gradient would be exactly the identity at every depth. That is the highway.
Veit et al. (2016) read the expansion literally: a residual network is an ensemble of $2^L$ paths, one for each subset of layers you pass through the skip versus the block. Most of the gradient comes from the short paths.
Delete a layer from a trained ResNet and it barely notices; delete a layer from a plain network and it collapses. In a transformer that is why you can drop blocks, add blocks, and swap in adapters without retraining from scratch: every block is a correction to a stream that works without it.
Why does the stream need a norm at all?
The skip solves the gradient problem and creates a scale problem. If every block adds something to the stream, the stream grows. With $L$ blocks each contributing a unit-variance update, $\lVert h_L \rVert$ grows like $\sqrt{L}$ if the updates are uncorrelated and like $L$ if they are aligned.
Either way the input to block 80 is much larger than the input to block 1, and a block that was tuned for unit-scale inputs, including the attention scores of Part 2 that need $O(1)$ logits, is now seeing something else.
Normalization is the fix: before a sublayer reads from the stream, rescale what it reads to a standard size. LayerNorm does it per token, across the $d$ features:
\[\text{LN}(x) = \frac{x - \mu}{\sqrt{\sigma^2 + \epsilon}} \odot \gamma + \beta, \qquad \mu = \frac{1}{d}\sum_i x_i, \quad \sigma^2 = \frac{1}{d}\sum_i (x_i - \mu)^2\]Every token becomes a vector with zero mean and unit variance across its features, then a learned per-feature gain $\gamma$ and bias $\beta$ let the network undo it where it wants to. Note what it does not do: it never looks at another token, and it never looks at another example in the batch.
The 30-second version
BatchNorm normalizes each feature using statistics across the batch, which is fine for images and terrible for sequences. Sequences have variable length and padding, so the batch statistics at position 200 are computed from whichever examples happen to be that long. Per-device batches in distributed training are small, which makes the statistics noisy. And at inference an autoregressive model processes one token at a time with batch size one, so BatchNorm needs running averages that no longer match training. LayerNorm uses only the token's own $d$ features: no batch dependence, identical at train and test time, works for any length and any batch size.Pre-norm or post-norm: where does the norm go?
Now the question with the interesting history. The 2017 paper placed the norm after the residual add:
\[h_{l+1} = \text{LN}\big(h_l + f_l(h_l)\big) \qquad \text{(Post-LN)}\]This is the layout in the famous figure. And it has a problem that took a couple of years to name: the norm sits on the residual path. The identity highway from the previous section goes through a LayerNorm at every layer, and LayerNorm is not the identity.
Its Jacobian rescales the gradient by $1/\sigma$ and projects out the mean direction, so the clean $I$ term is gone and you are back to a product of $L$ matrices, just better-behaved ones than in a plain stack.
Xiong et al. (2020) worked out what this does at initialization: in Post-LN the gradient with respect to the parameters of the last layers is large and independent of depth, whereas in Pre-LN it shrinks like $1/\sqrt{L}$. Liu et al. (2020) add that in deep Post-LN models the gradient reaching the early layers can vanish outright.
Large top-layer gradients at step zero are why Post-LN transformers need learning rate warmup: a full-size step at the start of training blows the model up. Warmup is not a hyperparameter preference; it is a patch for the norm placement.
The alternative moves the norm off the highway and into the branch:
\[h_{l+1} = h_l + f_l\big(\text{LN}(h_l)\big) \qquad \text{(Pre-LN)}\]The stream is never normalized. Each sublayer reads a normalized copy, computes its update, and adds it to the untouched stream. The identity term in the gradient is exactly the identity again, gradients are well-scaled at every depth from step zero, and Xiong et al. showed you can drop warmup entirely. GPT-2 shipped with this layout in 2019 and essentially every model since has kept it.
The toy above is a stack of random MLP sublayers with exact reverse-mode gradients, four ways. At a gain of exactly 1.0 the plain stack is flat, which is the knife’s edge every careful initialization scheme tries to balance on; nudge the gain to 1.2 or 0.8 and the gradient explodes or vanishes exponentially with depth.
The residual stack without any norm explodes in both panels: each block adds a unit-scale update to a stream that is already unit scale, so activations and gradients both compound. That is the problem GPT-2’s $1/\sqrt{2L}$ initialization scaling addresses. The two normed stacks do not care about the gain or the depth: push the slider to 96 layers and the Pre-LN gradient stays within a factor of a few of one everywhere.
In this caricature the Post-LN gradient is flat too, because a LayerNorm cancels an MLP branch’s gain exactly; a real transformer’s attention sublayer at initialization is not so well matched, which is what Xiong et al. worked out.
The bottom panel is the price. In Pre-LN the stream norm grows with depth, because nothing ever normalizes it. Each block’s update is unit-scale, so relative to a stream that has grown to size $\sqrt{l}$, block $l$’s contribution is worth $1/\sqrt{l}$ of what block 1’s was.
Later layers matter less, not because they learned less but because the sum they are adding to is already large. Liu et al. measured this and found that a well-tuned Post-LN model, when you can get it to train, slightly beats a Pre-LN one of the same size, precisely because its later layers pull their weight.
The 30-second version
Pre-LN keeps the residual path clean, so gradients are well-scaled at every depth from initialization: no warmup needed, trains stably at 80 or 100 layers. Post-LN puts the norm on the residual path, which makes the gradient scale at initialization depend on depth, large at the top layers and unreliable at the bottom, so it needs warmup and careful learning rates; but a Post-LN model that does train tends to use its later layers more effectively. Pre-LN's cost is that the unnormalized stream grows with depth, so each successive block's contribution is relatively smaller. Every major LLM uses Pre-LN; the research that is still active is about recovering Post-LN's per-layer effectiveness without its instability.That active research is worth knowing by name. DeepNet (Wang et al., 2022) kept Post-LN and made it stable to a thousand layers by scaling the residual branch and its initialization with depth. Gemma 2 uses both norms, one before the sublayer and one after its output, which is sometimes called sandwich norm.
And the most direct attack on the growing-stream problem is to stop adding block outputs with weight one and let the model choose how much of each previous layer to read, which is the attention-residuals idea I wrote about in a previous post. All of these are answers to the same question: how do you keep the highway and still make layer 80 count?
There is also a cheaper trick that GPT-2 introduced and everyone copied: scale the output projection of each residual branch ($W_O$ in attention, $W_2$ in the FFN) by $1/\sqrt{2L}$ at initialization. With $2L$ branches each adding something of variance $1/(2L)$, the stream stays at unit variance at initialization regardless of depth. It does not stop the growth during training, but it means the model starts from a sane place.
What is RMSNorm, and why did it replace LayerNorm?
One last norm question, because every Llama-style model answers it the same way.
The 30-second version
RMSNorm drops the mean subtraction and the bias: $\text{RMSNorm}(x) = x / \sqrt{\frac{1}{d}\sum_i x_i^2 + \epsilon} \odot \gamma$. It keeps the property that matters, invariance to the scale of the input, and gives up re-centering, which turns out to contribute almost nothing. One fewer reduction over $d$, no bias, and a measurable speedup on a layer that is memory-bound rather than compute-bound. Same quality in every comparison anyone has published, so it became the default.Zhang and Sennrich (2019) proposed it with exactly that argument: they hypothesized the benefit of LayerNorm was re-scaling rather than re-centering, tested it, and found no loss in quality with 7 to 64 percent less time in the normalization layer. Llama, Mistral, Gemma, DeepSeek, and Qwen all use it. The formula is short enough to write from memory in an interview and you should be able to.
Rapid fire: can you do these from memory?
- Expand $\prod (I + J_i)$ and point at the term that makes deep networks trainable.
- Why did a 56-layer plain network underfit compared to a 20-layer one?
- Write LayerNorm. Which axis does it normalize over, and why not the batch axis?
- Where does the norm sit in Post-LN, and what does that do to the residual path?
- Why does Post-LN need warmup and Pre-LN does not?
- What grows with depth in Pre-LN, and what is the consequence for later layers?
- Write RMSNorm and name what it drops from LayerNorm.
- Why scale $W_O$ and $W_2$ by $1/\sqrt{2L}$ at initialization?
Part 4 is about a symmetry that attention has and language does not: permute the tokens and attention does not notice. Positional encoding is how you break it, and there are three ways that keep coming up in interviews.