Inside attention · “the cat sat down”
A worked example, one operation at a timeFour illustrative features, one attention head. These are teaching numbers, not trained embeddings. Values are rounded for display; calculations use full precision. Play, pause, or scrub to inspect each operation.
Open video ↗Read the walkthrough & assumptions
Words become rows. Here each word is one token. An embedding lookup assigns “the” [1, 0, 1, 0], “cat” [0, 1, 1, 0], “sat” [1, 1, 0, 1], and “down” [0, 1, 0, 1]. Stacking these gives a 4 × 4 input matrix H. Rows are tokens; columns are features. Real embeddings have many more dimensions and are learned during training.
Three views of the input. Normalize each row to form X. This example uses RMS normalization with unit gain and zero epsilon. Projection matrices produce Q = XWQ, K = XWK, and V = XWV, mapping four features to two. Follow “sat”: each query coordinate comes from its input row dotted with one column of WQ.
Attention chooses a mixture. Dot the query for “sat” with every key and divide by √2. Causal masking replaces the score for the future word “down” with negative infinity. Softmax converts scores into nonnegative weights summing to one. Multiply each value vector by its weight and add: this is the attention context for “sat”. Queries and keys determine weights; values provide content.
Scope. This film opens up one attention head, not the entire transformer block. It omits positional encoding, other heads, output projection, the residual connection, and the feed-forward network. Toy weights illustrate arithmetic, not learned linguistic relationships. The animation is silent and does not start automatically.