Knowledge gradient ascent / visual studiesMixture of experts · new direction

One FFN becomes a bank.
Only a few run.

Mixture of Experts changes the feed-forward sublayer inside a transformer block. Keep attention fixed; increase the available FFN weights without evaluating them all for every token.

0:00

Swipe the diagram horizontally, or view in landscape.

Behind the motion

A visual teaching model, not DeepSeek-V3. Token capsules stand for full vectors; internal circuits show schematic activation, not literal neuron measurements. All routing and output numbers are computed.

Previous arithmetic walkthrough ↗
Model details, limitations and an accessible walkthrough

Dense: all four tokens pass through the same feed-forward network. Sparse: a small router scores the expert bank for each token, selects two networks, sends the same vector to both, and combines their outputs. Experts are learned feed-forward networks, not preassigned subject specialists. The network pictures are schematic. The top diagram shows a pre-norm transformer block; its attention and residual connections do not change. The detailed view remains inside the feed-forward slot throughout. Replacing one dense FFN with a bank is an architectural change, not a weight-preserving decomposition of a trained network.

In this toy, hidden vectors for “the cat sat down” are [2,1], [−1,2], [1,2], [−2,−1]. The first four router columns are [1,0], [0,1], [−1,0], [0,−1]. Routing uses a deterministic top-2 selection and softmax restricted to those scores. There are no shared experts, capacity drops, balancing losses or routing noise. Row-vector notation is used.

For “sat”, scores [1,2,−1,−2] select E2 and E1 with weights 0.731 and 0.269. Their outputs [2,2] and [0,4] combine to [1.46,2.54]. Every toy expert is a bias-free 2→3→2 ReLU network, with 12 weights. Four experts plus the router store 56 weights; eight store 112. Two experts run in either case, but the wider router increases active weights from 32 to 40. The comparison dense FFN has hidden width 3N and 12N weights: all 48 or 96 weights run per token. This matches the bank’s expert-weight capacity, excluding its additional router. The benefit is more learned FFN capacity for a given expert-compute budget, or less expert computation at comparable capacity. More stored parameters do not guarantee better quality.

Changing the bank size may change the selected experts; top-2 still runs two. The additional toy experts use different router columns and reuse the four example FFN weight sets. The source contains the exact toy weights and rendering logic. The outer block shows attention, normalization and both residual connections for context; only the feed-forward computation is expanded and numerically illustrated. Memory, routing and communication remain real costs. Background: Sparsely-Gated Mixture-of-Experts.