What does the FFN add that attention cannot?
For a long time I thought of the feed-forward network as the part of the transformer you mention for completeness. Attention was the idea. The FFN was two linear layers with a ReLU between them, the most boring object in deep learning, sitting there because every block needs a nonlinearity somewhere.
Then I counted the parameters and it was two thirds of the block. Then I read the interpretability work and it turned out that this is where the facts live. Then Mixture of Experts made it the one component everybody is scaling.
The boring part is the part the model thinks with, and the interview question about it has a precise answer that most people fumble.
The 30-second version
Attention's output for a token is a convex combination of value vectors: nonlinear in the weights, but linear in the values. It can move information between positions but cannot apply an arbitrary nonlinear function to a token's own representation. The FFN is the only place in the block that does: an elementwise nonlinearity between two linear maps, applied to each token independently. Structurally, attention mixes along the sequence axis with weights computed from the data; the FFN mixes along the feature axis with weights learned once and shared across all positions. Without it, stacked attention layers collapse toward rank one and the model can route but not compute.Why is attention linear in the values?
Write out what a single attention head returns for token $i$:
\[o_i = \sum_j \alpha_{ij}\, v_j = \sum_j \alpha_{ij}\, W_V x_j\]The weights $\alpha_{ij}$ come from a softmax, which is nonlinear. But look at where the $x_j$ enter: linearly, through $W_V$, and then in a weighted sum. Fix the attention pattern and the output is a linear function of the inputs. The only nonlinearity is in deciding how much of each token to take, never in what is done to it.
Two things follow. First, an attention layer cannot compute something like “square this token’s third feature” or “if feature 5 is large then activate feature 9.” Those need an elementwise nonlinearity applied to the token, and attention has none.
Second, and this is the result that made me take the FFN seriously, a stack of pure attention layers actively destroys information. Dong et al. (2021) proved that without skip connections and MLPs, the output of a deep attention stack converges to a rank-one matrix doubly exponentially in depth: every token’s representation collapses toward the same vector.
Averaging averages of averages produces a constant. The skip connection slows the collapse and the FFN is what counteracts it, by pushing each token back out into its own direction.
Which axis does each sublayer mix along?
The cleanest way to hold this is to look at the residual stream $H \in \mathbb{R}^{T \times d}$ as a grid with two axes, and notice that the two sublayers each mix along exactly one.
Attention mixes along $T$. Row $i$ of the output is a weighted sum of rows of $V$, and every column stays in its own column: feature 3 of the output only ever sees feature 3 of the values. The weights are data-dependent, computed fresh for every input.
The FFN mixes along $d$. Row $i$ of the output depends only on row $i$ of the input, and every feature of that row is combined with every other feature through $W_1$ and $W_2$. The weights are fixed, learned once, and identical for every position and every sequence.
Neither can do the other’s job. A transformer is the alternation: gather from context, process alone, gather, process.
MLP-Mixer made this explicit by replacing attention with a fixed linear mix along $T$ and keeping the FFN, and it works, worse than attention but far better than nothing, which tells you the FFN was carrying more of the load than its reputation suggested.
Why is the FFN a key-value memory?
The mechanistic view that changed how I think about it is from Geva et al. (2021). Write the FFN as
\[\text{FFN}(x) = W_2\, \sigma(W_1 x) = \sum_{i=1}^{d_{ff}} \sigma(k_i \cdot x)\, v_i\]where $k_i$ is the $i$-th row of $W_1$ and $v_i$ is the $i$-th column of $W_2$. Read it as attention.
Each $k_i$ is a key: $k_i \cdot x$ asks how well the input matches pattern $i$. The nonlinearity $\sigma$ keeps the matches and zeros the rest. Each $v_i$ is a value that gets added to the residual stream in proportion to its match. The FFN is a lookup over $d_{ff}$ memory slots, with keys and values that are learned rather than computed from context, and a ReLU where attention would have a softmax.
Geva et al. found that in a trained language model the keys in lower layers match shallow patterns (specific n-grams, surface forms) and the keys in upper layers match semantic ones (text about a topic, sentences that end a certain way), and that the corresponding values push the output distribution toward tokens that plausibly follow.
The follow-up work on model editing, ROME among others, located factual associations like “the Eiffel Tower is in Paris” in mid-layer FFN weights and rewrote them by changing a single $v_i$. When people say the model “knows” something, this is the layer they mean.
Why is the hidden width four times the model width?
The 30-second version
It is a convention from the 2017 paper ($d = 512$, $d_{ff} = 2048$) that has survived because nothing beats it by enough to bother. In the memory view, $d_{ff}$ is the number of key-value slots, so wider means more stored patterns per layer. Scaling-law studies found the loss is nearly flat for $d_{ff}/d$ anywhere from about 1 to 10 at fixed parameter count, so 4 is a safe middle. It gives the FFN $8d^2$ parameters against attention's $4d^2$, two thirds of the block. Gated activations like SwiGLU use three matrices instead of two and shrink the hidden width to about $8d/3$ to keep the parameter count the same.The honest answer is that 4 is a historical accident with a wide basin of good values around it. Kaplan et al. (2020) varied the ratio across more than an order of magnitude and saw the loss change by a few percent.
What is not an accident is that the FFN gets most of the parameters: a slot-based memory wants to be wide, and there is no equivalent pressure on the attention projections, which only need to be wide enough to route.
Why did the nonlinearity become a gate?
The 2017 FFN used ReLU. BERT and GPT-2 switched to GELU, a smooth version. Since PaLM and Llama the standard is a gated unit:
\[\text{FFN}_{\text{SwiGLU}}(x) = W_2\, \big(\text{Swish}(W_1 x) \odot W_3 x\big), \qquad \text{Swish}(z) = z\,\sigma(z)\]Three matrices. One branch is squashed through an activation, the other is a plain linear map, and they are multiplied elementwise before the output projection. Under the memory view, the gate $\text{Swish}(W_1 x)$ decides which slots fire and $W_3 x$ decides what magnitude they fire with, a separation the two-matrix version cannot make.
Shazeer (2020) tested the variants, found the gated ones consistently better at equal parameter count, and declined to explain why, attributing it “as all else, to divine benevolence.” Nobody has improved much on that explanation.
Multiplicative interactions between two linear projections of the same input are strictly more expressive than a single activation, and the model uses the extra freedom.
To keep the parameters at $8d^2$ with three matrices, Llama sets $d_{ff} \approx \frac{2}{3} \cdot 4d = \frac{8d}{3}$, rounded up to a multiple of 256. Llama 3 8B has $d = 4096$ and $d_{ff} = 14336$.
Why does Mixture of Experts replace the FFN and not attention?
If the FFN is a memory of $d_{ff}$ slots, the obvious way to give the model more memory without more compute is to have several FFNs and only consult a few per token.
\[\text{MoE}(x) = \sum_{e \in \text{TopK}(g(x))} g_e(x)\, \text{FFN}_e(x), \qquad g(x) = \text{softmax}(W_r x)\]A small router $W_r$ scores each of $E$ expert FFNs for the token, the top $k$ are run, and their outputs are combined with the router weights. The parameters scale with $E$; the compute per token scales with $k$.
Mixtral 8x7B has eight experts of which two run per token, for 47B total parameters and 13B active. DeepSeek-V3 has 256 routed experts plus one always-on shared expert, runs eight routed per token, and reaches 671B total with 37B active.
The 30-second version
Two reasons. The FFN is where the parameters are, two thirds of the block, so sparsifying it is where the savings are. And the FFN is position-wise: each token is processed independently, so routing different tokens to different experts is natural and the experts never need to see each other's tokens. Attention is the layer where tokens interact, so an expert that only saw some tokens would break the mechanism. Modern MoE keeps one dense attention layer per block and swaps the FFN for a bank of experts, which fits the memory view exactly: more slots, consulted selectively.The hard part of MoE is not the math, it is load balancing. A router left to itself sends most tokens to a few favourite experts, which starves the others and wastes the capacity.
The Switch Transformer added an auxiliary loss pushing the routing toward uniform; DeepSeek-V3 replaced the loss with a per-expert bias that is nudged up when an expert is underused, which balances load without distorting the gradient. Expect a question on this if the interviewer has trained one.
What does the universal approximation proof say the FFN is for?
One more thing worth having in your pocket. Yun et al. (2020) proved that transformers are universal approximators of continuous sequence-to-sequence functions with compact support, and the proof uses the two sublayers for two distinct jobs: attention to build a contextual mapping in which every token’s representation becomes unique given its context, and the FFN to then map each of those unique representations to whatever output is required.
Attention alone cannot do the second step. The division of labour in the proof is the division of labour in the block.
Rapid fire: can you do these from memory?
- Write the attention output for one token and point at where it is linear.
- What happens to a deep stack of attention layers with no FFN and no skip connections?
- Which axis of the $(T, d)$ matrix does each sublayer mix along, and which has data-dependent weights?
- Rewrite $W_2 \sigma(W_1 x)$ as a sum over slots and name the keys and values.
- Why is $d_{ff} = 4d$, and how flat is the loss around it?
- Write SwiGLU. Why is the hidden width $8d/3$ rather than $4d$?
- Why is MoE applied to the FFN and not to attention?
- Name the failure mode of a naive MoE router and one fix.
Part 6 closes the series with the diff: everything that has changed between the 2017 block and the one inside a 2026 model, and the one-line reason for each.