Why does a transformer need positional encoding at all?

Take the sentence “dog bites man.” Shuffle it to “man bites dog.” Feed both through a stack of attention layers with no positional information and look at the row for “bites.” It is identical in both cases. Not similar: identical, to the last bit.

This part is mostly pictures. Each one is something an interviewer can hand you a marker for.

They will ask Why does a transformer need positional encoding at all?
The 30-second version Self-attention is permutation equivariant. If $P$ is a permutation matrix and you feed $PX$ instead of $X$, the queries, keys and values are all permuted the same way, the score matrix becomes $P(QK^\top)P^\top$, the row-wise softmax commutes with that, and the output is $P \cdot \text{Attn}(X)$: the same rows in a different order. The FFN acts on each row separately, so it does not break the symmetry either. Without position information the model sees a bag of tokens. Positional encoding injects the order, either by adding a position-dependent vector to the input or by modifying the attention scores as a function of the offset between tokens.

Why is attention a set operation?

The whole proof is one line. Permute the rows of $X$ by $P$, and $Q$, $K$, $V$ are permuted the same way:

\[\text{softmax}\big(PQK^\top P^\top\big)\,PV = P\,\text{softmax}\big(QK^\top\big)\,V\]

The rows come out permuted and nothing else changes. One asterisk for later: this assumes every token can see every other. A causal mask is not symmetric, and the last picture on this page is about that loophole.

How does the 2017 sinusoidal encoding work?

The original paper added a fixed vector to each input, indexed by position:

\[PE(pos, 2i) = \sin\!\left(\frac{pos}{10000^{2i/d}}\right), \qquad PE(pos, 2i+1) = \cos\!\left(\frac{pos}{10000^{2i/d}}\right)\]

Do not read the formula yet. Step through the figure: its seven steps are the whole argument, from why the two obvious encodings fail to what attention gets out of a row of clocks. It carries one five-token sentence throughout, “the cat sat on the”, where the same word sits at positions 0 and 4 and only the position vector can tell the two copies apart.

Step 6 is the property the paper cared about. Written down:

\[PE(pos + k) = R_k \, PE(pos), \qquad R_k = \bigoplus_i \begin{pmatrix} \cos k\omega_i & \sin k\omega_i \\ -\sin k\omega_i & \cos k\omega_i \end{pmatrix}\]

A fixed rotation per pair, depending only on $k$. That was the paper’s hope: a linear layer can learn “$k$ tokens back” as a rotation.

BERT and GPT-2 replaced the table with a learned one of shape $(L_{\max}, d)$. Same idea, no structure: position 2049 does not exist in a table of 2048, and nothing tells the model that 10 is next to 11.

They will ask Sinusoidal versus learned positional embeddings: what is the tradeoff?
The 30-second version Both are absolute and both are added once at the input. Learned embeddings are a free table, cost $L_{\max} \times d$ parameters, and cannot represent positions beyond the table. Sinusoidal embeddings are parameter-free, defined for every position, and have the structural property that shifting by $k$ is a fixed rotation, which is a hint toward relative offsets. In-distribution they perform about the same, which is why the 2017 paper reported no difference. Neither generalizes well to lengths beyond training, because the model has never seen those vectors or those distances. Modern models use neither: they put position into the attention scores directly.

Where does the position go in?

Both absolute schemes add the position once, at the bottom of the stack. What a head actually wants is an offset: “the token three back,” not “index 847.”

Shaw et al. (2018) put a learned vector per offset into the key, T5 reduced it to a scalar bias per offset bucket, and ALiBi dropped the parameters entirely and subtracts $m \cdot (i - j)$ from the score with a fixed slope per head. All of them share the picture on the right. RoPE is the version that won.

How does RoPE turn position into rotation?

Rotary position embedding (Su et al., 2021) keeps the sinusoidal frequencies and changes one thing: instead of adding a vector to the input, it rotates the query and the key inside every attention layer.

\[q_m = R(m\theta)\, q, \qquad k_n = R(n\theta)\, k, \qquad q_m \cdot k_n = q^\top R(m\theta)^\top R(n\theta)\, k = q^\top R\big((n - m)\theta\big)\, k\]

Shift $m$ and $n$ together: every hand moves, the wedge does not, the score does not. The rotations are absolute, one per token, which is exactly what a KV cache needs: rotate each key once by its own position, store it, never touch it again. The “all pairs” tab stacks the sinusoidal frequencies. Fast pairs resolve neighbours, slow pairs track the long range, and the summed score fades with distance before any training has happened.

They will ask How does RoPE work, and why has it replaced the alternatives?
The 30-second version RoPE rotates each 2-D pair of query and key dimensions by an angle proportional to the token's position, with a different frequency per pair. Because $R(m\theta)^\top R(n\theta) = R((n-m)\theta)$, the dot product depends only on the relative offset $n - m$. It has no parameters, it is applied inside every attention layer rather than once at the input, it is compatible with the KV cache because each key is rotated once by its absolute position, and it has a built-in decay with distance. Learned embeddings cannot extrapolate; sinusoidal ones fade through the layers; ALiBi is relative but has no content-dependent structure. RoPE gets relative position, per-layer injection, and cacheability at once, and it has a knob, the base, that can be scaled to extend context.
Key idea Rotate queries and keys by their absolute position and the dot product sees only the difference. RoPE is relative position implemented with absolute operations, which is why it works with a cache and at every layer.

What breaks when you stretch the context?

The follow-up that always comes: trained at 4k, wanted at 32k. What breaks?

The slow pairs break: they never completed a turn in training, so past the training length they point at angles the weights have never seen. Every extension trick is a way of choosing $\theta_i$ so that the new angles land on old ones. Position interpolation squeezes every position and blurs the fast pairs.

NTK-aware scaling raises the base, which stretches the slow pairs and leaves the fast ones alone. YaRN does it per pair. Llama 3 raised the base from 10 000 to 500 000 and then extended from 8k to 128k in stages of continued pretraining.

Can a model know its position with no encoding at all?

Back to the asterisk. The permutation proof assumed full attention.

With a causal mask, query $i$ sees exactly $i + 1$ keys. A head that attends uniformly returns their mean, and the spread of a mean over $i + 1$ items shrinks like $1/\sqrt{i+1}$. A later layer can read $i$ off that. Haviv et al. (2022) trained causal language models with no positional encoding at all and they nearly match models with explicit positions; Kazemnejad et al. (2023) found they extrapolate at least as well as several explicit schemes on synthetic tasks.

Nobody ships it, because RoPE is free and the leaked signal is noisy. But it sharpens the answer to the opening question: causal attention is not quite a set operation, and that asymmetry is enough for a model to bootstrap a coordinate system on its own.

Rapid fire: can you do these from memory?

  1. Prove that self-attention is permutation equivariant in one line. What does the causal mask do to the proof?
  2. Write the sinusoidal encoding. Why sines and cosines rather than the integer position?
  3. Show that $PE(pos + k)$ is a linear function of $PE(pos)$.
  4. Name two things learned absolute embeddings cannot do.
  5. Derive $q_m \cdot k_n = q^\top R((n-m)\theta) k$ from the rotation identities.
  6. Why is RoPE compatible with a KV cache, and why is it applied in every layer?
  7. Which RoPE frequencies break when you exceed the training length, and what do position interpolation and YaRN do about it?
  8. How can a causal model with no positional encoding know where a token is?

Part 5 is about the component I underrated the longest: the feed-forward network, which holds two thirds of the parameters and does something attention structurally cannot.