Inside Mixture of Experts · “the cat sat down”
A worked example, one operation at a timeA worked example: four experts, top-2 routing, two hidden features. Illustrative weights, not a trained model or DeepSeek-V3 simulation. Displayed numbers are rounded; calculations use full precision.
Open video ↗Read the walkthrough & assumptions
Input. Start after attention and normalization, at the feed-forward sublayer. Our illustrative hidden vectors for “the cat sat down” are [2, 0], [0, 2], [1, 2], and [2, 1]. They are not raw word embeddings. Follow the row for “sat”: x = [1, 2].
Router. Using row-vector notation, s = xWᵣ. The router matrix has rows [1, 0, 1, −1] and [0, 1, 1, 0], giving scores [1, 2, 3, −1]. Select experts 3 and 2. Softmax over just those selected scores gives weights 0.731 and 0.269; the other weights are zero. This deterministic toy omits routing noise, capacity limits, shared experts and load balancing. DeepSeek-V3 uses a different routing scheme, explored in the detailed walkthrough below this figure.
Experts. Each is a separate two-layer feed-forward network with its own weights: fᵢ(x) = ReLU(xWᵤₚ,ᵢ)W𝒹ₒ𝓌ₙ,ᵢ. For expert 2, the up-projection has rows [1, 1, −1] and [0, −1, 1]; its output [1, −1, 1] becomes [1, 0, 1] after ReLU. The down-projection has rows [1, 0], [0, 1], [1, 2], yielding [2, 2]. Expert 3's up-projection has rows [0, 1, −1] and [1, −1, 0]; its down-projection has rows [0, 2], [1, 0], [1, 1], yielding [0, 4]. Real experts are much wider and often use gated activations such as SwiGLU.
Combine. y = 0.269[2, 2] + 0.731[0, 4] ≈ [0.54, 3.46]. Both selected experts receive the same input vector; their outputs are weighted and added. The residual addition outside the MoE sublayer is not shown.
Capacity versus work. For N equal-sized experts with Pₑ weights each and a router with Pᵣ weights, this sublayer stores NPₑ + Pᵣ weights but uses kPₑ + Pᵣ per token. Here each bias-free expert has 2 × 3 + 3 × 2 = 12 weights, and the router has 2 × 4 = 8: 56 weights stored, 32 used per token. This is not an end-to-end speedup estimate. Attention, routing, communication and storage still have costs. Experts are not assigned human-defined topics.
Background: The Sparsely-Gated Mixture-of-Experts Layer. This film uses a simplified, noise-free routing example.