A toy stack of random MLP sublayers at initialization, run four ways. Top: how much gradient reaches each layer's input. Bottom: the size of the hidden state itself.
Width d = 48, f(x) = W2 ReLU(W1 x) with He initialization scaled by the gain, loss = u·hL for a fixed unit vector u. Gradients with respect to hl by exact reverse mode. Try gain 1.2: the plain stack explodes, the un-normalized residual stack explodes in both panels, and the two normed stacks do not care. A caricature of a transformer, not a simulation of one.