What does an MoE cost to train and to serve?

Parts 2 and 3 were about the router as a piece of mathematics. This part is about what happens when you actually have to run it, and the answer is that a Mixture of Experts layer stops being a matrix multiply and becomes a distributed systems problem.

The reason is simple to state. A dense layer’s weights are the same for every token, so you can shard them however you like and the tokens never move. An MoE layer’s weights are chosen per token, the chosen ones are on other devices, and so the tokens move instead.

That single change swaps an all-reduce for two all-to-alls, replaces static message sizes with sizes the router decides at runtime, and, during generation, quietly undoes most of the arithmetic-intensity argument that made sparsity attractive in the first place.

The GPU series drew expert parallelism as one axis on the parallelism map. This part is what happens along that axis.

They will ask Why is serving a Mixture of Experts harder than serving a dense model with the same active parameter count?
The 30-second version Three reasons, in increasing order of how often people miss them. Memory: the whole bank has to be resident because routing is decided per token mid-forward-pass, so you size HBM by total parameters and get throughput from active ones. Communication: the experts a token wants are on other GPUs, so every layer runs two all-to-alls whose payload sizes are unknown until the router has run, which fights static shapes and CUDA graphs. And arithmetic intensity, which is the one that decides throughput. Each decoding token picks $k$ of $N$ experts independently, so a batch of $B$ tokens touches about $N(1 - (1 - k/N)^B)$ distinct experts per layer. At batch one that is $k$. At batch 128 on DeepSeek-V3 it is 252 of 256. So the bytes you read stop scaling with $k$ and start scaling with $N$ while the FLOPs keep scaling with $k$, and the intensity settles at $N/k$ times worse than a dense model of the same weights. DeepSeek-V3 needs roughly thirty-two times the concurrent batch of a dense model to reach the same point on the roofline, which is exactly why the published serving setups run expert parallelism a hundred and forty-four ways during decode.

The figure computes all of that from DeepSeek-V3’s shapes and H100-class hardware, with every assumption on the screen so you can disagree with it.

Why is expert parallelism the natural split?

Because a bank of experts has a seam that a dense matrix does not. Expert parallelism sits alongside the other four axes in the Parallelism series; here it is the sparse model’s own bill.

Tensor parallelism cuts every matrix by rows or columns, so each device computes a partial result and the partials have to be summed: an all-reduce, every layer, with a message size known at compile time. That works, and for the attention half of an MoE model it is still what you do.

Expert parallelism gives whole experts to whole devices. Nothing needs summing across devices, because each expert’s output is complete where it was computed. What has to move instead is the tokens, out to their experts and back with the results.

For DeepSeek-V3 the arithmetic is stark. One MoE layer’s experts are 11.3 billion parameters. Spread 257 experts over 64 GPUs and each holds four of them, 177 million parameters per layer instead of 11.3 billion. That is not an optimisation, it is the only reason 671 billion parameters fit anywhere.

What travels, and how much of it?

Per token, per layer, in each direction: the hidden state out to each destination, and one result vector back per expert.

Dispatch sends $d$ values per destination. Combine brings back $d$ values per expert. DeepSeek’s DeepEP dispatches in FP8 and combines in BF16, so the return journey moves twice the bytes of the outward one. With $d = 7168$ and $k = 8$ that is about 57 kB out and 115 kB back per token per layer, before any optimisation.

Now multiply. Fifty-eight MoE layers, two collectives each in the forward pass and two more in the backward, is 232 all-to-alls per training step, each of which is a full every-GPU-to-every-GPU exchange with $P^2$ flows.

The lever that matters most is node-limited routing. Cap the number of distinct nodes a token is allowed to reach, send it across the slow network once per node rather than once per expert, and fan it out over NVLink inside. DeepSeek-V3 capped it at four nodes, which turns eight network destinations into at most four.

Is an MoE layer a FLOP problem or a network problem?

Do the arithmetic and it is not close.

Take the figure’s default: expert parallelism 64 across eight nodes, 8192 tokens per GPU, node-limited routing on, an H100 at 989 TFLOP/s with 450 GB/s of NVLink and a 400-gigabit port that moves 50 GB/s. Each GPU computes about 65,000 token-expert pairs, which is 6.5 TFLOP and 6.6 milliseconds at peak. The traffic it has to move is 705 MB across the network, which is 14.1 milliseconds, plus 1.4 GB over NVLink, which is 3.1.

Communication is 2.6 times the arithmetic. That is the whole story of why MoE systems work exists, and it is why an interviewer asking “what is the bottleneck in an MoE layer” is not asking about matrix multiplies.

Change any input and the ratio moves, which is why the sliders are there. What does not move is the shape of the answer. The arithmetic and the traffic both grow linearly with tokens, so batching does not help the ratio. What helps is the interconnect, the precision of the payload, and how many nodes a token is allowed to touch.

key idea An MoE layer's expert arithmetic and its all-to-all traffic both scale linearly with the number of tokens, so the ratio between them is a property of the architecture and the interconnect, not of the batch size. You cannot batch your way out of it. You can only overlap it, shrink the payload, or bound the fan-out.

How do you hide the all-to-all?

A ratio above one only hurts if the two things happen in sequence, and essentially all the systems work since GShard is about making them happen at the same time.

Pipeline schedules put a different micro-batch’s arithmetic underneath this one’s collective. DeepSeek’s DualPipe runs the pipeline in both directions at once so there is always compute travelling the other way, which the GPU series covered as the reason their 64-way expert parallelism was affordable at all.

Kernel libraries attack it from below. DeepEP’s dispatch and combine kernels are written for zero or minimal streaming-multiprocessor occupation, which matters more than it sounds: an all-to-all implemented as an ordinary kernel steals the multiprocessors the expert matrix multiplies want, so a communication kernel that uses no SMs is not just faster, it stops competing.

DeepSeek’s own April 2026 report puts a number on how far this goes. They split the experts into waves, so that the computation of the current wave, the token transfer for the next, and the result-sending of the completed ones all proceed at once.

And they state the balance condition directly: because a token-expert pair costs $6hd$ FLOPs but only $3h$ bytes, each gigabyte per second of interconnect bandwidth suffices to hide the communication for 6.1 teraflop per second of compute. Above that threshold, more bandwidth buys nothing. They report 1.50 to 1.73 times speedups on general inference and up to 1.96 on latency-sensitive work, and they shipped the fused kernel as MegaMoE.

That result is also why V4 removed the node-limited routing constraint that V3 needed. Once the communication is genuinely hidden, the cap that bounded it stops earning its cost in routing quality.

Why is decoding a sparse model harder than decoding a dense one?

This is the part that surprises people, and it is the single best thing to have on your fingertips in this whole series.

Training moves thousands of tokens together. Generation moves one per sequence per step. Each decoding token picks its $k$ experts independently, so with $B$ tokens decoding together the expected number of distinct experts touched in a layer is

\[\mathbb{E}[\text{experts touched}] = N \left( 1 - \left(1 - \tfrac{k}{N}\right)^{B} \right)\]

For DeepSeek-V3, $N = 256$ and $k = 8$. At $B = 1$ that is 8 experts, 3% of the bank. At $B = 32$ it is 163. At $B = 128$ it is 252, which is 98% of the bank.

So the bytes you have to read stop scaling with $k$ and start scaling with $N$, while the FLOPs keep scaling with $k$. Arithmetic intensity is the ratio of exactly those two things, and it collapses.

At $B = 128$, one MoE layer reads about 11 GB of expert weights to do about 100 gigaflops: an intensity of 9 FLOP per byte, against the H100 ridge of 295 from Part 2 of the GPU series. A dense model holding the same weights would be at 256 at that batch, because it reads its weights once and every token uses all of them.

As $B$ grows the ratio settles at exactly $N/k$, which for V3 is 32. A sparse layer needs thirty-two times the concurrent batch of a dense one to sit at the same point on the roofline.

What does that do to how you serve it?

It splits the machine in two, which is roughly what every 2026 inference stack converged on.

Prefill is compute-bound. The matrix multiplies are already large, so the all-to-all is pure overhead and narrow expert parallelism wins.

Decode is bandwidth-bound. Spreading the experts over more devices shrinks the weights each one has to read and frees the HBM that a larger batch needs, so wide expert parallelism wins.

DeepSeek published their own serving configuration and it is exactly this split: expert parallelism 32 across four nodes for prefilling, and expert parallelism 144 across eighteen nodes for decoding, with two routed experts and one shared expert per GPU. Their stated reason is the arithmetic above, in their words: only 8 of 256 experts are activated per layer, so the system needs an extremely large overall batch to give each expert a batch size worth having.

The two phases want opposite settings of the same knob. That is the systems argument for running them on separate machines, and it is the answer to “how would you deploy this” that shows you have thought past the model card.

The other levers are all about the skew from Part 3, which does not go away at inference. The same DeepSeek writeup runs 32 redundant routed experts in both phases and an expert-parallel load balancer whose objective is to minimise the maximum dispatch load across GPUs, because the all-to-all is synchronous and the slowest rank sets the step time for everyone.

Why is temperature zero not determinism?

Because on a capacity-limited serving stack your token’s routing depends on the other tokens in the batch.

If expert capacity is enforced, an expert that fills up drops whatever arrives next, and what arrives next depends on which other requests happened to be batched with yours. Two identical prompts at temperature zero can take different paths through the model because someone else’s traffic changed at the same moment.

The dropless kernels remove that particular version of the problem, but a subtler one remains: routing is a top-$k$ over floating-point scores, so any numerical difference at all, a different kernel, a different batch shape, a different accumulation order, can flip which expert wins. Discrete selection turns a rounding difference into a completely different computation.

That is a good thing to raise unprompted in a systems interview, because it explains a class of production bug that looks impossible: the model gives different answers, the seed is fixed, and nothing in the request changed.

What is the memory floor, really?

Total parameters, always. There is no way around it, and this is the most common misconception about MoE in the whole subject.

The reason is timing. The router decides mid-forward-pass, per token, which experts are needed. There is no window in which to fetch a weight from host memory or from disk, so every expert has to be resident in HBM somewhere before the layer starts.

Offloading the cold experts to the host is a real technique and it works exactly where you would expect: at batch one, where a token touches eight of 256 experts, most of the bank genuinely is idle. As soon as you batch, the coverage formula above says you touch nearly all of them, and offloading turns into streaming the whole model over PCIe every step.

So the honest sentence is: an MoE saves FLOPs, not bytes. It is a way to buy capacity with memory instead of with arithmetic, and memory is the resource you have to be able to afford.

Rapid fire: can you do these from memory?

  1. Why does expert parallelism need an all-to-all where tensor parallelism needs an all-reduce?
  2. Compute the dispatch and combine bytes per token per layer, given $d$, $k$ and the precisions.
  3. What does node-limited routing bound, and what does it cost?
  4. Given tokens per GPU, expert parallel width and an interconnect, compare the traffic time against the expert arithmetic time.
  5. Derive the expected number of distinct experts touched by a batch of $B$ decoding tokens.
  6. Why does an MoE layer's arithmetic intensity settle at $N/k$ times worse than a dense layer's?
  7. Why do prefill and decode want opposite expert-parallel widths?
  8. Give two reasons a Mixture of Experts is not deterministic at temperature zero.
  9. Does an MoE let you serve a large model on less memory? Answer precisely.

Part 5 leaves the mechanism behind and asks what is actually unsettled: which models are sparse in September 2026, what changed in the last year, and the list of questions where a confident answer is a bad sign.