Where the gradient goes in a deep stack

A toy stack of random MLP sublayers at initialization, run four ways. Top: how much gradient reaches each layer's input. Bottom: the size of the hidden state itself.

plain: h ← f(h)residual, no norm: h ← h + f(h)Post-LN: h ← LN(h + f(h))Pre-LN: h ← h + f(LN(h))
depth L = 32 weight gain = 1.0

Width d = 48, f(x) = W2 ReLU(W1 x) with He initialization scaled by the gain, loss = u·hL for a fixed unit vector u. Gradients with respect to hl by exact reverse mode. Try gain 1.2: the plain stack explodes, the un-normalized residual stack explodes in both panels, and the two normed stacks do not care. A caricature of a transformer, not a simulation of one.