What is unsolved about Mixture of Experts in 2026?

The first four parts described a mechanism the field has largely converged on. A bank of narrow experts, sigmoid affinities, top-8, one shared expert or none, balanced by a bias rather than a loss, sharded by expert parallelism, dropless.

The interesting thing about September 2026 is that the convergence has started to come apart, and the lab that broke it first is the one that set most of it. DeepSeek-V4, published in April, changed the gate, removed the routing constraint that V3 needed, replaced the dense layers at the bottom of the stack with hash-routed experts, and moved the expert weights to four bits.

This part is the state of play, dated, with the questions where a confident answer is a bad sign. I have tried to keep the distinction visible throughout between what a lab measured, what a lab asserted, and what nobody has settled.

They will ask What is the current state of the art in Mixture of Experts, and what would you be worried about if you had to build one?
The 30-second version The 2025 recipe was fine-grained experts, sigmoid affinities, top-8, auxiliary-loss-free balancing with a per-expert bias, dropless routing, and expert parallelism with a fused all-to-all. Sparsity ratios went from about 4 with Mixtral to about 30 with Kimi K2 and DeepSeek-V4, while active parameters stayed between roughly 13 and 50 billion, so the growth all went into total parameters. What would worry me is four things. Hyperparameters do not transfer cleanly across sparsity: the September 2026 work on 1,800 pretraining runs finds optimal learning rate and batch size shift with the activation ratio in a way neither parameter count explains. Post-training is unstable: routing is a discrete function of numerics, so under reinforcement learning about 10% of activated experts flip after each gradient update on a 30B model, which breaks importance sampling. Expert specialisation is contested, and appears to depend on whether the model was upcycled. And the balancing mechanisms both have regimes where they misbehave, which is why DeepSeek ship a bias controller and a small auxiliary loss at the same time.

The figure is the landscape: eleven open-weight models with published numbers, the sparsity trend, what routing alternatives exist, and the upcycling question.

Which models are sparse, and how sparse?

Every open-weight frontier release since 2024 is a Mixture of Experts, and their configs are public, so this is one of the few places in the subject where you can be exact.

Mixtral 8x7B, in December 2023, stored 47 billion parameters and used 13. DeepSeek-V3, a year later, stored 671 and used 37. Meta’s Llama 4 Maverick stores 400 and uses 17, with 128 experts and an unusual routing scheme: each token goes to one routed expert plus a shared one, in layers that alternate with dense ones.

Qwen3-235B-A22B has 128 experts, top-8, no shared expert, across 94 layers. Kimi K2’s config lists 384 routed experts plus one shared. GLM-4.5 lists 160 plus one. OpenAI’s gpt-oss-120b has 128 experts with top-4 routing and the expert weights quantised to MXFP4.

DeepSeek-V4, in April 2026, comes in two sizes: V4-Pro at 1.6 trillion total and 49 billion active, and V4-Flash at 284 billion and 13.

Two trends fall out of that list. The sparsity ratio went from 3.6 to more than 30 in under three years. And the active-parameter axis barely moved: the frontier settled somewhere between 13 and 50 billion active parameters and spent all its growth on the other axis.

That second observation is the one worth being able to state, because it says what the binding constraint actually is. Compute per token is set by the inference budget and by what a serving fleet can afford. Total parameters are set by how much HBM you can put behind it. Sparsity is the only knob that moves capacity without moving the first number.

What changed in DeepSeek-V4?

More than the parameter count, and the specific changes are the best available snapshot of where the recipe is going.

The affinity function moved. V3 computed expert scores with a sigmoid; V4 changed it to the square root of a softplus. That is a small edit to one line and it is the first time a frontier lab has publicly moved off the gate everyone had copied.

Node-limited routing is gone. V3 capped each token at four destination nodes to bound the network fan-out, which was the lever Part 4 spent a section on.

V4 removes the constraint, and the reason is in their own systems section: with the wave-based overlap they describe, the communication within a single MoE layer takes less time than the computation, so once it is fused into one pipeline the compute is the bottleneck and the system tolerates lower interconnect bandwidth. Their threshold is that each gigabyte per second of interconnect suffices to hide 6.1 teraflop per second of compute, and past it more bandwidth buys nothing.

The dense layers at the bottom are gone too. V3 kept plain feed-forward networks in its first three blocks. V4 replaces them with MoE layers using hash routing, where a predefined hash of the token id decides the experts. No router, no gradient, no balancing problem, in exactly the layers where the earlier literature found routing looks most like a partition of the vocabulary anyway.

And the routed expert parameters are stored in FP4. Balancing is still auxiliary-loss-free with a small sequence-wise loss alongside, which is the one part of the V3 recipe that survived untouched.

key idea The constraints an architecture is designed around are the ones the systems layer has not yet solved. DeepSeek-V3 capped routing at four nodes because the all-to-all was exposed; V4 removed the cap because a fused kernel hid it. Read every architectural restriction as a statement about the hardware of its year.

Has routing moved past top-k?

Not in what ships, and that is itself the interesting answer.

The alternatives are all well known. Expert choice inverts the argmax: instead of every token picking $k$ experts, every expert picks its best $c$ tokens. Load is then balanced by construction, with no auxiliary loss and no bias controller at all. The catch is that an expert’s choice depends on the other tokens in the batch, including later ones, which in a causal decoder leaks information backwards in time, and that a token can end up chosen by nobody.

AI2 ran the comparison directly for OLMoE and reported that dropless token choice beat expert choice on every task at the same token budget, while expert choice ran about 20% faster per device. That is the trade in one line: expert choice buys throughput and balance and pays in quality and in causality.

Hash routing goes the other way and does not learn at all. The Hash Layers work in 2021 reported that routing tokens by a fixed hash was competitive with learned routing, which is an uncomfortable result if you believe the router is doing something clever. It sat as a curiosity for five years and then reappeared in DeepSeek-V4’s first blocks.

So the honest summary is that top-$k$ token choice is still the default, that nothing has beaten it convincingly at scale, and that the two most credible alternatives are one that cannot be used in a decoder and one that does not learn.

Should you upcycle a dense checkpoint?

You already have a good dense model. Copy its feed-forward layer $N$ times, perturb the copies, bolt on a router, and keep training. That is sparse upcycling, and it is the cheapest way to get an MoE.

Google’s paper on it put the crossover at roughly 120% of the original dense training budget: below that, upcycle, above it, start over.

AI2 tried it for OLMoE and got a different number. They upcycled their OLMo-1B checkpoint after 2 trillion tokens and trained for 610 billion more. An otherwise equivalent MoE trained from scratch caught up after 500 billion extra tokens and started beating it around 600 billion. That is 25% of the original budget, not 120%, and they rejected upcycling for their final model.

Both are right about their own setting, and the useful thing is the reason they differ. Beyond the compute bracket, AI2’s stated objection was that an upcycled model inherits hyperparameters tuned for the dense model it came from, and the previous section says those do not transfer. There is a second reason, and it connects to the specialisation question below: experts that all start from the same weights start correlated, and it takes training to pull them apart.

The practical reading is that upcycling is a way to spend a small budget well, not a way to reach a frontier model.

Do the hyperparameters transfer?

This is the newest result on the list, dated 8 September 2026, and it is the one I would raise unprompted.

Muon, muP and the whole hyperparameter-transfer literature rest on the idea that you can tune a small proxy and scale the settings up along a known law. The paper “Hyperparameter Scaling Laws Across MoE Sparsity” tests that across sparsity levels with 1,800 pretraining runs, six activated-parameter scales, models up to 6 billion total non-embedding parameters, about 20 trillion tokens, and roughly 200,000 H800-hour equivalents.

The finding is that conventional hyperparameter scaling laws are insufficient for ultra-sparse models. Optimal learning rate and optimal batch size both shift with the activation ratio, and those shifts cannot be explained by either total or activated parameter count alone. Their resolution is that at fixed sparsity the optimal batch size is a power law in the token count and the optimal learning rate a power law in compute, and that across sparsity levels the activation ratio enters both as an additional multiplicative factor.

The interview version of this: if you tune on a dense proxy and transfer to a model activating a sixty-fourth of its parameters, you should expect the transfer to be wrong, and neither of the two parameter counts you might reach for will tell you by how much.

Do experts specialise?

The popular claim is that a Mixture of Experts learns a maths expert, a code expert, a biology expert. The evidence does not support that, and being able to say why is a genuinely differentiating answer.

Mistral looked for it in Mixtral and did not find it. Their own routing analysis says the assignment distribution is very similar across arXiv papers, PubMed abstracts and philosophy papers at every layer. What they did find is structure of a different kind: Python’s self and indentation tokens route consistently, and consecutive tokens pick the same first-choice expert about 28% of the time at layer 15 against a 12.5% random baseline. That is syntax and temporal locality, not subject matter.

ST-MoE found clear specialisation in their encoder experts, on punctuation and verbs and proper nouns, and explicitly did not find it in the decoder. They also looked for language specialisation in a multilingual model and reported that experts handled English, Japanese, French and Chinese indiscriminately.

OLMoE, trained from scratch rather than upcycled, does report domain and vocabulary specialisation, with one layer-0 expert nearly 100% specialised to arXiv, and their hypothesis for the disagreement is that Mixtral was upcycled from Mistral, so its experts all started from the same optimum and had less room to diverge.

There is a mechanism underneath that makes the whole thing less mysterious. The balancing scope from Part 3 acts directly against specialisation: if you compute the dispatch fractions over a micro-batch of a few sequences, you are demanding that a batch of pure code spread itself over the whole bank.

Computing them over the global batch instead improves both perplexity and measured domain specialisation, at scales up to 42.8 billion parameters. Some of what looks like a fact about Mixture of Experts is a fact about a reduction in a framework.

So: token class and vocabulary, yes. Subject matter, contested and initialisation-dependent. Language, no.

What breaks when you reinforcement-learn a sparse model?

The routing, and it is the sharpest modern failure mode.

Top-$k$ is a discrete function of continuous scores, so any numerical difference at all can flip which expert wins. In reinforcement learning that matters more than anywhere else, because the algorithms depend on comparing the probability a token had under the old policy against its probability under the new one, and if the two policies ran different experts you are comparing two different networks.

The Qwen team put a number on it while introducing GSPO. On the 48-layer Qwen3-30B-A3B base model, after each gradient update and for the same rollout sample, roughly 10% of the activated experts differ between the new policy and the old one, and the effect becomes more prominent in deeper models.

Their workaround was routing replay: cache the experts the old policy activated and replay those routing decisions when computing the importance ratios. They call it essential for normal convergence of GRPO on MoE models, and they note it costs memory, communication, and some of the model’s effective capacity.

GSPO’s own contribution is to sidestep it by working at the sequence level rather than the token level, so it is not sensitive to individual token likelihoods and does not need the replay at all.

If someone tells you their reinforcement learning run diverges on a sparse model and is fine on a dense one of the same size, this is the first thing to check, and it is not a bug in their code.

Does sparsity change what else you have to change?

Yes, and the combinations are where the 2026 work is.

The obvious pairing is sparse attention. Parts 1 and 4 established that sparsity lives entirely in the feed-forward half of the block: all of DeepSeek-V3’s attention runs on every token, and the KV cache is untouched by any of it.

So a model that is sparse in the feed-forward layer and dense in attention has moved its bottleneck rather than removed it. That is why DeepSeek pairs MoE with multi-head latent attention to shrink the cache and, from V3.2 onward, with a sparse attention mechanism to cut the score computation.

Their natively trainable sparse attention work in February 2025 laid the groundwork. V4’s hybrid attention combines a compressed sparse mechanism with a heavily compressed one, and they report V4-Pro needing 27% of V3.2’s single-token inference FLOPs and 10% of its KV cache at million-token context.

The other pairing is quantisation, and it interacts with sparsity in a way dense models do not have to think about. Experts see wildly different amounts of data. A hot expert is calibrated on a diverse stream; a cold one is under-represented in whatever calibration set you use, and a shared expert sees everything.

Quantising them all identically is the obvious thing and the wrong one. gpt-oss quantises the MoE weights to MXFP4 and leaves attention, the router and the embeddings alone, which is the shape of the answer: the router is tiny and precision-sensitive, the experts are enormous and are where the bits are.

What would I actually say if asked what comes next?

One prediction and one refusal.

The prediction: the active-parameter axis stays roughly where it is and the total-parameter axis keeps climbing. The binding constraints are inference FLOPs and HBM bandwidth, and sparsity is the only knob that moves capacity without moving either. Nothing in the public sweeps has found a loss-based ceiling on how sparse a model can be; the ceilings that have actually been hit are systems ceilings and numerical ones.

The refusal: I would not predict what the router looks like. In eighteen months the field has kept top-$k$ token choice, moved from softmax to sigmoid, moved from an auxiliary loss to a bias controller, moved from sigmoid to the square root of a softplus, and put hash routing back into production.

Every one of those changes was small and none was forced by a theory. The honest summary is that the router is the least understood component of the most successful architecture of the decade.

That is also the reason the questions in Part 6 exist. The mechanism is learnable in an afternoon. Knowing which of its parts are load-bearing and which are conventions nobody has re-examined is the thing that takes longer.

Rapid fire: can you do these from memory?

  1. Name four open-weight models from the last two years with their total and active parameter counts.
  2. What did DeepSeek-V4 change from V3, and what does the removal of node-limited routing tell you about the systems layer?
  3. Why can expert-choice routing not be used unmodified in a causal decoder?
  4. State both published crossover points for upcycling versus training from scratch, and one reason they differ.
  5. Why does hyperparameter transfer from a dense proxy break on an ultra-sparse model?
  6. Summarise the evidence on expert specialisation at three different granularities.
  7. Why does reinforcement learning destabilise a Mixture of Experts more than a dense model, and what is routing replay?
  8. Why does a sparse feed-forward layer not help the KV cache, and what do the labs pair it with?

Part 6 is the self-test: three levels of questions with answer sketches, built from everything above.