Why would you make a model sparse?
The Transformer series ended with a block I could write from memory. The GPU series put a clock on it. Then an interviewer asked me something that sits across both: “DeepSeek-V3 is a 671 billion parameter model that trains for about the price of a 37 billion parameter one. How?”
The answer is one structural change, applied to one sublayer, and every consequence of it fills five parts. In a dense transformer, parameters and arithmetic per token are the same quantity. You cannot buy one without the other. A Mixture of Experts cuts that link.
Compare two versions of the same feed-forward sublayer. On the left, every token runs one wide FFN. On the right, a router selects two of four smaller FFNs and combines their outputs. Follow the hidden vector for “sat” in “the cat sat down” to see which weights are evaluated and how the mixing equation is built. This is a small teaching example with top-2 softmax routing and ReLU experts, not a simulation of DeepSeek-V3.
Explore the full DeepSeek-V3 block and parameter counters
The detailed walkthrough puts one token through one layer of DeepSeek-V3 and writes the working as it goes: normalise, attend, add it back, normalise again, and then the single box a Mixture of Experts replaces. Step 2 opens that box in a dense layer, step 3 opens it in a sparse one, and the last step puts the two next to each other.
Watch the counters at the bottom of the first step, because they answer the question everyone asks next. The arithmetic per token is the same on both tracks, and that is deliberate: nine experts of 2,048 is 18,432, the dense width exactly. What differs is the pool it comes from. The dense token multiplies the same 396 million parameters every time. The MoE token multiplies a different 396 million, chosen out of 11.3 billion.
What you get is capacity that costs storage but not FLOPs. What you pay is a routing decision on every token in every layer, a load-balancing problem that has no clean solution, and a network bill that turns a layer of matrix multiplies into a distributed systems problem.
This part is the trade in its simplest form: what the layer becomes, where the parameters go, and what the two columns of the ledger actually say.
The 30-second version
In a dense transformer the feed-forward network is one wide matrix triple, and every token multiplies against all of it, so FLOPs per token are about twice the parameter count. An MoE layer replaces that one network with a bank of narrower ones, the experts, plus a small router that scores every expert for the token and picks the top few. Only the chosen experts run. Total parameters are set by the size of the bank; arithmetic per token is set by how many you pick. DeepSeek-V3 stores 671 billion parameters and touches 37 billion per token, so it trains and serves on a 37B compute budget with a 671B capacity. That is the whole argument: at fixed compute you get more parameters, and more parameters is what scaling laws want. What you pay is memory, because every expert has to be resident somewhere; communication, because the experts a token wants are on other GPUs; and a routing problem, because nothing in the loss makes the router use the bank evenly.Why are parameters and FLOPs the same knob in a dense model?
Take the feed-forward sublayer of a dense block. With SwiGLU it is three matrices of shape $d \times d_{ff}$, and a token passes through all three. The parameter count is $3 d\, d_{ff}$ and the arithmetic is one multiply and one add per parameter, so the FLOPs are $6 d\, d_{ff}$.
Twice the parameters, exactly. That relationship holds for every dense matrix in the model, which is where the familiar $6ND$ estimate for training compute comes from: six FLOPs per parameter per token, forward and backward.
So in a dense model there is one dial. Want more capacity? Pay more arithmetic on every token of every batch of the entire training run, and again on every token you ever serve. The scaling laws say capacity is what you want. The compute budget says you cannot have it.
Everything about Mixture of Experts follows from noticing that the token does not actually need all of that width.
What does an MoE layer replace the feed-forward network with?
A bank of $N$ smaller feed-forward networks, called experts, and a router. The router is a single matrix of shape $d \times N$. It turns the token into $N$ scores, the top $k$ are selected, those $k$ experts run on the token, and their outputs are summed with the router’s scores as weights.
\[y = \sum_{i \in \text{TopK}(s)} g_i \cdot \text{FFN}_i(x), \qquad s = W_r x\]That is the entire idea. Total parameters scale with $N$. Arithmetic per token scales with $k$. Two dials where a dense model has one.
The experts are usually much narrower than the layer they replace. DeepSeek-V3’s dense layers use $d_{ff} = 18432$; its experts use $d_{ff} = 2048$, a ninth as wide, and it keeps 256 of them plus one that always runs. That choice has a name, fine-grained expert segmentation, and Part 2 is about why it is the right one.
Which layers get a bank, then? Not all of them. The report is blunt about it: “We substitute all FFNs except for the first three layers with MoE layers.” So the first three blocks are ordinary transformer blocks, the other fifty-eight carry the experts, and attention is left alone in every one of the sixty-one.
That is the box the expandable DeepSeek-V3 walkthrough above stops on, and the two steps after it open the box both ways: as one wide network in a dense layer, and as a bank of narrow ones everywhere else.
Where do DeepSeek-V3’s 671 billion parameters actually go?
Two counts, one model. Total parameters count the weights stored; active parameters count the weights used for one token. Start with one MoE layer below: choose a word and watch the selection change. The unused experts stay in the bank.
The two sliders separate the trade-off: a bigger bank stores more weights; a bigger selection evaluates more weights. At a fixed selection size, adding experts only increases the small router’s contribution to the active count. Across all 58 MoE layers and the rest of the model, the published totals are about 671B stored and 37B active per token. Open the breakdown above to see where they come from.
What does the sparsity actually buy?
Compare at fixed compute, which is the comparison that matters, because compute is the budget. A dense model with DeepSeek-V3’s arithmetic per token has 37 billion parameters, by the identity from two sections ago. V3 has eighteen times that.
The claim that this is worth something is empirical, and the papers are specific. Google’s Switch Transformer reported up to a sevenfold speedup to a fixed quality against dense T5 baselines at equal compute, and about fourfold against T5-XXL. DeepSeek’s own MoE paper reported a 145B sparse model matching their dense 67B while doing a fraction of its arithmetic.
The intuition is the one from Part 5 of the Transformer series: the feed-forward layer behaves like a key-value memory. More slots is more storage. A given token only needs to look up a few of them, and looking up the ones it does not need was never doing any work.
What sparsity does not buy is anything on the attention side. All 11.4 billion parameters of DeepSeek-V3’s attention run on every token. Sparsity lives entirely in the feed-forward half of the block.
What does it cost?
Three bills, and the interview is usually about the second and third.
Memory first, because it is the one people get wrong. An MoE saves FLOPs, not bytes. Every expert has to be resident in HBM somewhere, because routing is decided per token in the middle of a forward pass and there is no time to fetch a weight from anywhere slower.
How many bytes that is depends on the format the weights are stored in. DeepSeek publish V3 only in FP8, with a conversion script for anyone who wants BF16, so a byte a parameter is the right count for serving it: about 671 GB, against the 640 GB that eight H100s hold.
Say that carefully in an interview, because the training claim is not the same claim. V3 was trained in FP8 mixed precision, not in FP8 throughout. The expensive GEMMs run in FP8. The embedding, the output head, the MoE gating, the normalisations and the attention operators are held at BF16 or FP32 on purpose, and the master weights, the gradients and the optimizer states are all kept higher.
Most checkpoints still ship in BF16, two bytes a parameter, which would put the same model at 1.34 TB. That single choice of format is the largest lever on this line, which is why the last step of the expandable DeepSeek-V3 walkthrough puts a toggle on it.
Training is the worse number either way. At the sixteen bytes per parameter the GPU series counted for Adam in mixed precision, 671 billion of them come to 10.7 TB before a single activation. V3’s own choices shave that, keeping optimizer states in BF16 and caching activations in FP8, and it is still a number no single node can hold.
Communication second. The experts a token wants live on other devices, so an MoE layer sends every token across the network and brings the results back, twice per layer, in a collective whose message sizes nobody knows until the router has run. Part 4 does that arithmetic.
Balance third. Nothing in the training objective cares whether the experts get used evenly, and a router left alone will not use them evenly. That is Part 3, and it is the failure mode that quietly deletes the capacity you paid for.
Why the feed-forward layer and not attention?
Two reasons, and having both is the difference between a memorised answer and an understood one.
The feed-forward layer is where the parameters are. In a dense block it is about two thirds of the weights, so it is the only place where sparsifying is worth the machinery.
And it is position-wise. Each token passes through the feed-forward layer independently of every other token, so sending different tokens to different experts changes nothing about what the layer means. Attention is the opposite: it is the one place where tokens interact, and an expert that only saw some of the tokens would be computing a different function, not a cheaper one.
There is research on sparsifying attention heads, and there is sparse attention, which is a different idea about which score entries to compute rather than which weights to use. Neither is what “MoE” means when an interviewer says it.
Is a sparse model just a small model wearing a big number?
This is the sharp version of the question and it deserves a real answer rather than a defensive one.
At fixed pretraining loss, the sparse model is often the worse one. There is measured evidence that at equal perplexity a sparse model underperforms a dense one on reasoning-heavy downstream tasks, which is what you would expect if reasoning consumes inference compute that a sparse model does not spend.
But nobody chooses at fixed loss. They choose at fixed compute, or fixed cost per served token, and on both of those axes the sparse model wins, which is why every frontier open-weight release since 2024 has been one.
So the honest framing is that the sparse model is not a 671B model with a 37B price tag. It is a 37B-compute model with 671B worth of storage behind it, and the reason that is a good trade is that storage got cheaper faster than arithmetic did.
Rapid fire: can you do these from memory?
- Why are parameters and FLOPs per token the same number, up to a factor of two, in a dense layer?
- Write the MoE layer's forward equation, including where the gate value multiplies.
- Given layers, hidden size, expert width, expert count and top-k, compute total and active parameters.
- Does an MoE reduce the memory needed to serve a model? Say exactly which quantity you mean.
- Which parts of a transformer block stay dense in every shipped MoE, and why?
- Name the three bills that come with sparsity, and which part of the model each lands on.
- At fixed pretraining loss, is a sparse model better or worse than a dense one? At fixed compute?
Part 2 opens the router: what it computes, what the gate value is, and why every large model now has hundreds of small experts instead of eight big ones.