What changed between the 2017 block and a 2026 model?

Hand someone the 2017 paper and ask them to draw the block inside Llama 3. They will get the skeleton right, two sublayers on a residual stream, and almost every detail wrong. No encoder. No post-norm.

No LayerNorm at all, no sinusoidal table, no ReLU, no biases, no dropout, and the attention has eight times fewer key and value heads than query heads. The architecture is famously “unchanged since 2017,” and that is true only at the resolution of a block diagram.

This part is the diff. For each change: what it was, what it is now, and the one-line reason, because the reason is what the interviewer wants. Most of these reasons were derived in Parts 2 through 5; this is where they become a list you can recite.

They will ask What are the main differences between the original transformer and a modern LLM like Llama 3 or DeepSeek-V3?
The 30-second version Decoder-only instead of encoder-decoder. Pre-norm instead of post-norm, and RMSNorm instead of LayerNorm. RoPE applied to queries and keys in every layer instead of a sinusoidal table added at the input. Grouped-query attention, or DeepSeek's multi-head latent attention, to shrink the KV cache. SwiGLU instead of ReLU in the FFN, and in the biggest models a Mixture of Experts in its place. No biases, no dropout, much larger vocabularies. Under the hood, FlashAttention computes exactly the same attention without materializing the score matrix, and contexts have gone from 512 to 128k and beyond through RoPE scaling and local or sparse attention layers. The block diagram is the same; nearly every box in it has been swapped.

Why one stack instead of an encoder and a decoder?

The 2017 model translated, so it had an encoder to read the source and a decoder to write the target, joined by cross-attention. GPT kept only the decoder and trained it on one objective, predict the next token, and it turned out that a long enough prefix does the encoder’s job.

The source document is just the start of the sequence. One stack is simpler to scale, every token in the corpus is a training example, and the same model serves every task through the prompt. Encoder-only models still exist for embeddings and classification; encoder-decoder still makes sense when the input is fully known and very different in kind from the output. For “an LLM,” decoder-only is the answer.

What happened to the norms?

Part 3 covered the first two. The norm moved from the residual path into the branch (Pre-LN), because the identity path needs to stay clean for gradients to reach early layers without warmup. LayerNorm became RMSNorm, because re-centering was not doing anything and dropping it saves a reduction and a bias.

The third move is newer. Part 2 ended with attention logits growing during training until the softmax saturates mid-run and the loss spikes. QK-norm applies an RMSNorm to the queries and keys of each head before the dot product, which bounds the logit by a learned scale times $\cos\theta$ regardless of how large the weights get.

Gemma 3, OLMo 2, and most open models released since have used it. Gemma 2 also adds a norm on the output of each sublayer before it is added to the stream, the “sandwich” layout, and caps the attention and output logits with a $\tanh$ soft-cap. All of these exist because a run of trillions of tokens gives the weights time to drift somewhere the initialization never anticipated.

Why did position move into attention?

Part 4’s story in one line: the sinusoidal table was added to the input once and had to survive every layer, whereas RoPE rotates $q$ and $k$ inside every attention layer so that the score sees the relative offset directly. No parameters, cache-friendly, and a base you can raise to stretch the context. Llama 3 uses a base of 500000; longer contexts are reached with YaRN-style scaling and continued pretraining.

Why did attention get narrower on the key-value side?

This is the change with the best interview arithmetic.

They will ask What is grouped-query attention and why does every modern model use it?
The 30-second version During generation each new token attends over the cached keys and values of every previous token, and that cache is what limits batch size and context length in serving. Multi-query attention shares one K and V head across all query heads; grouped-query attention is the middle ground, one KV head per group of query heads. Llama 3 70B has 64 query heads and 8 KV heads, so the cache is 8 times smaller than with full multi-head attention, and the quality loss is within noise. The compute barely changes; what changes is the memory read per generated token, which is the actual bottleneck at inference.

Do the numbers for Llama 3 70B. Per token, the cache stores a key and a value for each layer and each KV head: $2 \times 80 \text{ layers} \times 8 \text{ heads} \times 128 \text{ dims} \times 2 \text{ bytes} = 320$ KiB. At 128k context that is 43 GB per sequence.

With 64 KV heads it would be 2.5 MiB per token and 344 GB per sequence, more than the model’s weights. Grouped-query attention, from Ainslie et al. (2023) building on Shazeer’s multi-query attention, is the difference between serving long contexts and not.

DeepSeek went further with multi-head latent attention. Instead of storing $K$ and $V$, store a single compressed latent vector per token (512 dimensions in DeepSeek-V3, plus a 64-dimensional decoupled component that carries the RoPE rotation) and reconstruct the keys and values from it on the fly.

The cache per token drops to about 70 KB across 61 layers, and the reconstruction matrices can be folded into the query and output projections so there is no extra cost at inference. It is a low-rank factorization of the KV cache with RoPE handled separately because a rotation does not commute with the compression.

What does FlashAttention change about attention?

They will ask What does FlashAttention change about attention?
The 30-second version Nothing about the math: it computes exactly $\text{softmax}(QK^\top / \sqrt{d_k})V$. What it changes is the order of operations so that the $T \times T$ score matrix is never written to GPU memory. It processes blocks of keys and values, keeps a running maximum and running sum for the softmax normalization, and rescales the partial output whenever the maximum changes. Attention had been bound by reading and writing that score matrix to HBM, not by FLOPs, so removing the traffic gives a large speedup and drops memory from $O(T^2)$ to $O(T)$. It is the reason long contexts are practical.

The trick worth being able to derive is the online softmax. For a row of scores processed in blocks, keep $m$ (the max so far), $\ell$ (the sum of $e^{s - m}$ so far), and $o$ (the partial output). When a new block arrives with max $m’$, set $m_{\text{new}} = \max(m, m’)$, rescale $\ell$ and $o$ by $e^{m - m_{\text{new}}}$, and add the new block’s contributions computed against $m_{\text{new}}$.

At the end $o / \ell$ is the exact attention output. Dao et al. (2022) built the kernel around this; FlashAttention-2 parallelized it better across the sequence, and FlashAttention-3 uses the asynchrony of Hopper GPUs. In PyTorch it is behind F.scaled_dot_product_attention.

How did the FFN get a gate, then go sparse?

Part 5. ReLU became SwiGLU, three matrices with hidden width about $8d/3$, because gated units train better at the same parameter count. In the largest models the dense FFN became a Mixture of Experts, because the FFN is where the parameters are and it is position-wise, so routing tokens to different experts is natural. DeepSeek-V3 activates 37B of 671B parameters per token this way.

What got deleted, and why?

Some of the biggest changes are subtractions. Biases on the linear layers are gone: they add memory traffic and buy nothing measurable at scale, and removing them helps with low-precision stability. Dropout is gone from pretraining: with one pass over trillions of tokens the model never sees a batch twice and underfitting is the failure mode, not overfitting.

Label smoothing is gone for the same reason. Tied input and output embeddings are a per-model choice now rather than a default. Vocabularies went the other way, from 37k to 128k or more, because a bigger vocabulary means fewer tokens per document, which is cheaper inference and more effective context.

How did contexts get a thousand times longer?

Contexts went from 512 to 128k, and in some models beyond a million. RoPE scaling handles the position side. The compute side is handled by not paying quadratic attention everywhere: Mistral and Gemma alternate sliding-window layers that only see a local span with full-attention layers, Gemma 3 at a ratio of five local to one global.

DeepSeek’s native sparse attention learns which blocks of keys to attend to. Ring attention and context parallelism split a single sequence’s attention across devices when it does not fit on one. None of these change the block; they change which entries of the score matrix are computed.

What changed beyond the block?

A few changes reach past the block diagram. PaLM and GPT-J computed the attention and FFN sublayers in parallel from the same normalized input rather than in sequence, for a throughput gain at large scale. DeepSeek-V3 trains with multi-token prediction, an auxiliary head that predicts the token after next, which densifies the training signal and enables speculative decoding.

The attention-residuals idea from Kimi, which I wrote about earlier, lets each layer choose how much of every previous layer’s output to read instead of summing them with weight one. And a growing set of models, Jamba, Qwen3-Next, and Kimi Linear among them, replace most of their attention layers with linear-time recurrent layers like Mamba or gated DeltaNet and keep full attention only every third or fourth block, because for most of the network the exact softmax lookup is not worth its cost.

Can you write the 2026 block in code?

Part 1 ended with the 2017 block in forty lines. Here is the same block with the diff applied: RMSNorm, RoPE, grouped-query attention, a fused attention kernel, SwiGLU, no biases.

import torch, torch.nn as nn, torch.nn.functional as F

class RMSNorm(nn.Module):
    def __init__(self, d, eps=1e-6):
        super().__init__()
        self.g, self.eps = nn.Parameter(torch.ones(d)), eps
    def forward(self, x):
        return x * torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + self.eps) * self.g

def rope(x, base=500_000):                        # x: (B, h, T, dk); rotate dim pairs
    T, dk = x.shape[-2], x.shape[-1]
    inv = base ** (-torch.arange(0, dk, 2, device=x.device) / dk)   # theta_i, (dk/2,)
    ang = torch.arange(T, device=x.device)[:, None] * inv           # m * theta_i, (T, dk/2)
    cos, sin = ang.cos(), ang.sin()
    x1, x2 = x[..., 0::2], x[..., 1::2]
    return torch.stack([x1 * cos - x2 * sin, x1 * sin + x2 * cos], -1).flatten(-2)

class LlamaBlock(nn.Module):
    def __init__(self, d, n_heads, n_kv_heads, d_ff):
        super().__init__()
        self.h, self.kv, self.dk = n_heads, n_kv_heads, d // n_heads
        self.norm1, self.norm2 = RMSNorm(d), RMSNorm(d)
        self.wq = nn.Linear(d, n_heads * self.dk, bias=False)
        self.wk = nn.Linear(d, n_kv_heads * self.dk, bias=False)     # fewer KV heads
        self.wv = nn.Linear(d, n_kv_heads * self.dk, bias=False)
        self.wo = nn.Linear(d, d, bias=False)
        self.w1 = nn.Linear(d, d_ff, bias=False)                     # SwiGLU: three matrices
        self.w3 = nn.Linear(d, d_ff, bias=False)
        self.w2 = nn.Linear(d_ff, d, bias=False)

    def attention(self, x):
        B, T, d = x.shape
        q = self.wq(x).view(B, T, self.h, self.dk).transpose(1, 2)    # (B, h, T, dk)
        k = self.wk(x).view(B, T, self.kv, self.dk).transpose(1, 2)   # (B, kv, T, dk)
        v = self.wv(x).view(B, T, self.kv, self.dk).transpose(1, 2)
        q, k = rope(q), rope(k)                                      # position goes in here
        k = k.repeat_interleave(self.h // self.kv, dim=1)            # GQA: share KV heads
        v = v.repeat_interleave(self.h // self.kv, dim=1)
        y = F.scaled_dot_product_attention(q, k, v, is_causal=True)  # fused, never stores TxT
        return self.wo(y.transpose(1, 2).reshape(B, T, d))

    def forward(self, x):
        x = x + self.attention(self.norm1(x))
        h = self.norm2(x)
        return x + self.w2(F.silu(self.w1(h)) * self.w3(h))         # SwiGLU

Llama 3 8B instantiates this with $d = 4096$, 32 heads, 8 KV heads, $d_{ff} = 14336$, 32 times. Every line that differs from Part 1’s version corresponds to one row of the diff above, and each row has a reason you can now give.

Where does this leave us?

Six parts ago the question was “walk me through a transformer.” The point of answering it at this length was never the list of components. It was to have the reasons on my fingertips, the way anyone who does this for a living should, so that under interview pressure the nuances do not slip.

The defence turned out to be five arguments. The residual stream is a bus that blocks nudge, which is why depth is cheap. Attention is a soft lookup whose softmax needs $O(1)$ logits to have a gradient, which is why the square root is there and why QK-norm came later.

Residuals give the gradient an identity path and pre-norm keeps that path clean. Attention is a set operation, so position has to be injected, and injecting it as a rotation inside the score is the version that survived. And the FFN is the only place a token’s own representation gets a nonlinearity, which makes it a memory, which makes it the place where parameters and sparsity go.

Everything in the 2026 diff follows from those five. That is the thing I want to walk into the room with: not the list, but the reasons, so that when the interviewer asks about a change I have not seen, I can work out why someone would make it.

And since an agent writes most of my blocks now and runs for hours doing it, the reasons are also what let me read what it wrote and know whether it is right.

Rapid fire: can you do these from memory?

  1. Why decoder-only? What does the prefix do that the encoder used to?
  2. Name the three norm changes and the reason for each.
  3. Compute the KV cache size per token for a model given layers, KV heads, head dim, and dtype.
  4. What does multi-head latent attention store, and why does RoPE need a separate component?
  5. Derive the online softmax update when a new block of scores arrives.
  6. Why were biases and dropout removed?
  7. How do local-global layer patterns and sparse attention change the cost of long context?
  8. Pick one row of the diff and explain it from a principle in Parts 2 to 5 rather than from memory.