Dense FFN versus Mixture of Experts: the same transformer sublayer, opened up
Explanation & comparison assumptions
Both diagrams occupy the feed-forward sublayer of a transformer. Attention, normalization and residual connections are not replaced. We follow “sat” in “the cat sat down” using an illustrative two-feature hidden vector x = [1, 2], not a raw word embedding.
The dense FFN expands two features to twelve, applies ReLU (negative values become zero), then projects back to two. Every token evaluates this whole FFN. Each edge in its diagram represents one scalar weight: 2 × 12 + 12 × 2 = 48 weights.
The MoE instead contains four smaller FFNs. Each expert has its own up- and down-projection matrices, with widths 2 → 3 → 2, so each has 12 weights. “Expert” does not mean a human-assigned subject area. Together the four FFNs store 48 weights, matching the dense network’s FFN weight capacity—not its outputs or quality.
A learned router matrix has rows [1, 0, −1, 0] and [0, 1, 0, −1]. Multiplying x by it gives scores [1, 2, −1, −2]. Top-2 selection chooses FFN₂ and FFN₁. Softmax over only those two scores produces mixing weights g₁ ≈ 0.269 and g₂ ≈ 0.731; both unselected weights are zero.
The same whole input vector enters both selected FFNs. They return FFN₁(x) = [0, 4] and FFN₂(x) = [2, 2]. Multiply each entire output by its router weight, then add matching features: y = g₁ FFN₁(x) + g₂ FFN₂(x) ≈ [1.462, 2.538]. The result has the same width as x and goes on to the usual residual addition. Intermediate displayed decimals are rounded; the calculation uses full precision.
Only two expert FFNs execute: 24 expert weights per token instead of 48. The router itself has 2 × 4 = 8 weights. Including it, the MoE stores 56 weights and evaluates 32 per token, versus the dense FFN’s 48 stored and 48 evaluated. These are weight counts, not a latency estimate. Routing, communication, memory traffic and other transformer layers also have costs.
This is a small, bias-free ReLU example, not a trained model or a DeepSeek-V3 routing simulation. Larger real experts often use gated activations. No shared experts, load-balancing losses, routing noise or capacity limits are simulated. Different tokens can select different pairs; extra capacity does not by itself guarantee better quality.