Can you answer these without looking anything up?

Five parts of prose is a good way to build an understanding and a bad way to test one. This part is the test.

Three levels. The first asks whether you can state the mechanism precisely: shapes, what is computed, what is stored, where the bytes and the FLOPs go. The second asks whether you can reason about the trade-offs, which is where most interviews actually live. The third is the frontier, where the useful answer is usually a question about what the interviewer means.

Every answer sketch is what a strong answer contains, not an essay. Nothing in them is a number this series has not already established, and if one of them surprises you, the part it came from is the one to reread.

They will ask Tell me about Mixture of Experts.
The 30-second version A prompt this open is asking you to choose the frame, so choose one and say why. The frame I would take: MoE is the one architectural change that separates parameter count from arithmetic per token, which lets you buy capacity with memory instead of with compute. Then the three consequences, in order: the router is a hard argmax so most of the bank receives no gradient from most tokens, which is why balancing is a mechanism rather than a loss term; the experts live on other devices so a layer becomes two all-to-alls with runtime-decided payloads; and at serving batch a batch touches nearly every expert per layer, so the arithmetic intensity is worse than a dense model's by the ratio of experts to active experts. Each of those has a number I can compute on the board if they want it.

Level 1: can you state the mechanism?

level 1the mechanism
  1. Mixtral 8x7B is eight seven-billion-parameter experts, and it has 47 billion parameters rather than 56. Where did the other nine billion go?

    answer sketch

    Nowhere: they were never there. Only the feed-forward sublayer is replicated across the eight experts. Attention, the norms, the embeddings and the output head exist once and are shared by every token regardless of routing, so counting eight copies of the whole 7B model double-counts everything outside the feed-forward layer. The name describes how the experts were sized, not how the parameters add up, which is exactly why the useful numbers are total and active rather than the name.

  2. A token arrives at an MoE layer as a vector of length $d$. Name every tensor computed on the way to the layer's output, with its shape.

    answer sketch

    Router scores, $N$ values, from one $d \times N$ matrix-vector product. A top-$k$ selection giving $k$ indices, which is not a tensor with a gradient. The $k$ gate values, scalars, from renormalising the chosen scores, possibly times a fixed scaling factor. Then $k$ expert outputs, each a $d$-vector, each produced by a full SwiGLU through a $d \times d_{ff}$ triple. Then a weighted sum of those $k$ vectors, plus the shared expert's $d$-vector if the model has one. The output is a $d$-vector, the same shape it arrived as.

  3. During the backward pass of an MoE layer, which quantities receive gradient and which receive none?

    answer sketch

    The $k$ chosen experts receive gradient through their outputs, scaled by their gate values. The router's weights receive gradient only through those $k$ gate values, because the top-$k$ selection is a hard argmax with no derivative. The experts that were not chosen receive nothing at all, and the router receives no information about whether one of them would have been better. That last sentence is the whole reason a Mixture of Experts needs a balancing mechanism that lives outside the loss.

  4. Given 61 layers of which the first 3 are dense, $d = 7168$, experts of width 2048, 256 routed plus 1 shared, top-8, and a 129k vocabulary with untied embeddings: compute total and active parameters.

    answer sketch

    One SwiGLU expert is $3 \times 7168 \times 2048$, about 44 million. A MoE layer holds 257 of those plus a $7168 \times 256$ router, so about 11.3 billion, of which 9 experts plus the router, about 398 million, are active. Fifty-eight MoE layers give 657 billion total and 23 billion active. Add the three dense feed-forward layers, the 61 attention blocks and the embeddings and you land on DeepSeek-V3's published 671 billion and roughly 37 billion. The point of doing it by hand once is that both columns come out of the same set of shapes.

  5. Does top-8 routing mean eight forward passes through the layer? Describe what a real kernel does instead.

    answer sketch

    No. You compute the routing for the whole batch, sort the token rows by expert index into one contiguous buffer, run a single grouped matrix multiply whose group boundaries are the expert offsets, then scatter the results back to token order and weight them by the gates. The two permutations do zero arithmetic and move gigabytes, so they sit at the far left of the roofline and run at memory bandwidth. That is where a naive implementation loses its throughput, and it is why block-sparse and grouped-GEMM kernels exist.

  6. Define expert capacity and the capacity factor, and say precisely what happens to a token that overflows.

    answer sketch

    Capacity is tokens per batch divided by the number of experts, times a capacity factor, and it is fixed before the routing is known so that the experts can run as batched matrix multiplies. A token arriving at a full expert is dropped: its computation is skipped and it passes to the next layer through the residual connection. It is not an error and nothing raises. Worth adding that the modern answer is to avoid the situation entirely, with block-sparse kernels that handle variable group sizes and remove the capacity factor as a hyperparameter.

  7. Write the auxiliary load-balancing loss, define both of its vectors, and give its value at perfect balance and at total collapse.

    answer sketch

    $L = \alpha N \sum_i f_i P_i$, where $f_i$ is the fraction of the batch's tokens dispatched to expert $i$ and $P_i$ is the mean router probability that expert received. Both sum to one over the experts, so the dot product is minimised when both are flat, and the factor of $N$ makes that minimum exactly 1. The maximum is $N$, when one expert takes everything. Switch Transformer used $\alpha = 0.01$ after sweeping five orders of magnitude around it.

  8. Where does DeepSeek's per-expert bias enter the computation, and where does it deliberately not?

    answer sketch

    It is added to the affinity score used for the top-$k$ comparison and nowhere else. The gate value that weights an expert's output is computed from the original affinity, without the bias. That separation is the whole design: the bias changes which experts are chosen without changing how much they count, so no term is added to the loss and no gradient is distorted. It is updated after each step by a small constant, down for over-subscribed experts and up for under-subscribed ones.

Level 2: can you reason about the trade-offs?

level 2the trade-offs
  1. At fixed active parameters, would you rather have 8 experts of width $W$ or 64 of width $W/8$? What changes and what does not?

    answer sketch

    Active parameters and FLOPs per token are identical by construction; that is the premise. What changes is the number of distinct expert teams a token can assemble, which goes up combinatorially, and that is the argument the fine-grained models were built on. What it costs is not arithmetic: a router matrix eight times larger, a longer sort in the kernel, and eight times as many network destinations per token in the all-to-all. So the answer is that it is nearly free in FLOPs and expensive in systems work, which is why it took a systems paper to make it practical.

  2. Does a Mixture of Experts reduce the memory needed to serve a model of a given quality?

    answer sketch

    No, and this is the most common misconception in the subject. Every expert must be resident in HBM because routing is decided per token in the middle of a forward pass and there is no window in which to fetch a weight from anywhere slower. Memory tracks total parameters; only speed tracks active ones. Offloading cold experts works at batch one, where a token really does touch a small fraction of the bank, and collapses under batching for the reason in the next question.

  3. A 256-expert top-8 model decodes 128 tokens together. How many experts does one layer read, and what does that do to arithmetic intensity?

    answer sketch

    Each token picks independently, so the expected number of distinct experts is $N(1 - (1-k/N)^B)$, which at $B = 128$ is about 252 of 256. The bytes read scale with $N$ while the FLOPs still scale with $k$, so intensity is worse than a dense model's holding the same weights by a factor that settles at $N/k$, which here is 32. Concretely the layer sits around 9 FLOP per byte against an H100 ridge of 295, so it is memory-bound by a wide margin and needs roughly thirty-two times the batch of a dense model to reach the same point.

  4. Your training run is dropping 15% of tokens. A colleague proposes doubling the capacity factor. Why is that the wrong first move?

    answer sketch

    Because dropping is a symptom of skew, and the capacity factor only decides how much skew you buffer. Doubling it makes the buffers half empty, so the grouped matrix multiply becomes mostly padding, and it hides the underlying problem rather than fixing it. The first move is to look at the per-layer load histogram and the maximum violation to find out how skewed the routing actually is, and then to fix the balancing: check that the auxiliary loss or bias controller is on, that it is applied per layer, and what batch the dispatch fractions are being reduced over.

  5. At step 40,000 your MoE's loss curve looks healthy. What could already be badly wrong, and how would you know?

    answer sketch

    Router collapse, which is invisible in the loss for a long time because a bank in which a few experts do all the work is still a functioning model, just a much smaller one than you are paying for. The instruments are the per-layer load histogram and the count of experts receiving zero tokens, both logged every step. The shape to watch for is that the imbalance sits flat for a long stretch and then accelerates, so a healthy-looking recent history is not reassurance. Collapse also usually starts in one layer, so a model-wide average will hide it.

  6. You profile a training step and the all-to-all is 2.5 times the expert compute. List the levers, ordered by how much they buy.

    answer sketch

    Overlap first, because it is the only lever that changes the sum into a maximum: a pipeline schedule that puts another micro-batch's arithmetic under this one's collective, and kernels that occupy no streaming multiprocessors so the communication stops competing with the compute for them. Then shrink the payload: dispatch in FP8 and combine in BF16 halves the outward leg. Then bound the fan-out with node-limited routing, so a token crosses the network once per destination node rather than once per expert. Raising the batch does not help, because the traffic and the arithmetic both scale with tokens.

  7. Why does top-1 routing with a renormalised gate leave the router with no gradient from the output, and what are the two ways out?

    answer sketch

    Renormalising over the chosen set makes the single gate identically 1, a constant, so its derivative with respect to the router's weights is zero and the router learns nothing from the task loss. It would be driven entirely by the balancing term. One way out is to keep the raw gate value as the multiplier rather than renormalising, which is what Switch Transformer did to make top-1 work. The other is to use $k \ge 2$, which is what the original formulation argued for and what essentially everything at scale now does.

  8. A colleague computes the balance loss per micro-batch instead of over the global batch, to save an all-reduce. What does that change about the model, not just the speed?

    answer sketch

    The formula is identical, which is why the difference is easy to miss, but the reduction scope changes what is being demanded. A micro-batch at frontier scale is a handful of sequences, so a per-micro-batch term effectively insists that every individual sequence spread itself evenly across all the experts, which forces a batch of pure code to use the whole bank and actively works against specialisation. Reducing over the global batch instead costs one all-reduce of an $N$-element vector, and published work at tens of billions of parameters reports better perplexity and more domain specialisation for it.

Level 3: can you defend a design decision?

level 3the frontier
  1. Argue against Mixture of Experts. You have a fixed compute budget per token and you may build either a sparse model or a dense one with the same active parameters.

    answer sketch

    The dense model fits in a fraction of the memory, so it serves on fewer devices and its economics are simpler. It has no all-to-all, so its layers have static shapes, it captures cleanly into a CUDA graph, it is deterministic, and a slow rank does not stall a synchronous collective. It has no routing to collapse, no balancing mechanism to tune and no expert placement to rebalance in production. And at equal pretraining loss there is evidence that sparse models transfer worse on reasoning-heavy tasks, plausibly because reasoning wants inference compute the sparse model does not spend. If the deployment is memory-constrained rather than compute-constrained, the dense model is simply the right answer.

  2. Shared expert or no shared expert? What would change your mind?

    answer sketch

    The argument for is that whatever every token needs does not have to be relearned redundantly in every routed expert, so isolating it frees the rest to differ, and three labs ship one. The argument against is combinatorial: making one expert always-on removes it from the choice set and, at small expert counts, deletes most of the combinations a token could assemble, which is what AI2 measured for OLMoE when the matched-compute ablation came out slightly against it. What would change my mind is granularity, because with hundreds of experts one always-on expert costs almost nothing in combinations, which may be why the large-bank models keep them and Qwen3 dropped theirs. It is a small effect and the labs genuinely disagree.

  3. Auxiliary loss or the per-expert bias controller? What does the evidence actually say?

    answer sketch

    DeepSeek's own numbers show the bias controller giving both better balance and slightly better perplexity, and the argument is that the auxiliary loss adds a gradient that is not trying to make the model better at the task. But the honest version has three caveats. The perplexity gap is small. DeepSeek ship both, with a small sequence-wise auxiliary loss alongside the controller and the controller's update rate set to zero for the last part of training. And the paper was rejected at ICLR 2025, with the authors conceding in the public reviews that the interference-gradient motivation was intuition rather than a demonstrated effect. So: probably the better default, on evidence weaker than its adoption suggests.

  4. A product manager says your MoE should have a medical expert and a legal expert. What do you tell them?

    answer sketch

    That the name is misleading and the published evidence does not support subject-matter specialisation emerging on its own. Mistral looked for it in Mixtral and found the routing distribution nearly identical for arXiv, PubMed and philosophy; what they found instead was syntax and positional locality. ST-MoE found token-class specialisation in an encoder and essentially none in a decoder, and no language specialisation at all. OLMoE, trained from scratch, does report domain specialisation, and hypothesises that upcycled models show less of it because their experts start identical. And the balancing scope matters: reducing the load term over a small micro-batch actively suppresses whatever specialisation would have emerged. If they want a medical expert, that is a routing constraint you would have to impose, not one you can expect.

  5. You have a strong dense checkpoint and budget for more training. Upcycle it into an MoE, or start over?

    answer sketch

    It depends on the size of the remaining budget relative to the dense run, and the two published crossovers are far apart: Google's sparse upcycling paper put it near 120% of the original budget, AI2 measured 25% for OLMoE and rejected upcycling on that basis. Beyond compute there are two structural costs to name. The upcycled model inherits hyperparameters tuned for the dense model, and those do not transfer cleanly across sparsity. And the experts all start from the same weights, so they start correlated, which is the mechanism OLMoE proposes for why upcycled models specialise less. My default: upcycle for a small budget, start over if you are going to train for a long time anyway.

  6. Reinforcement learning on your MoE diverges. The dense model of the same active size is fine. Diagnose it.

    answer sketch

    Start with the routing. Top-$k$ is discrete, so any numerical difference between the rollout engine and the training engine, or between the old and new policy, flips which experts run, and the importance ratio then compares two genuinely different networks. Qwen measured roughly 10% of activated experts changing after each gradient update on a 30B model, worse in deeper ones. The two known fixes are routing replay, caching the rollout's expert choices and replaying them when computing the ratios, which costs memory and some effective capacity; and moving the objective to the sequence level so it is not sensitive to individual token likelihoods, which is what GSPO does. I would confirm the diagnosis first by logging the fraction of tokens whose routing differs between the two engines.

  7. DeepSeek-V4 removed the node-limited routing that V3 needed. What has to be true for that to be a good idea, and would it be a good idea on your cluster?

    answer sketch

    It is a good idea exactly when the all-to-all is genuinely hidden under compute, because the cap only ever existed to bound exposed communication and it costs routing freedom. DeepSeek's own argument is that with a fused wave-based overlap the communication in a MoE layer takes less time than the computation, so the compute is the bottleneck and the system tolerates lower bandwidth; they give the threshold as a compute-per-bandwidth ratio. On a different cluster the answer flips as soon as your interconnect falls below that ratio or your kernels do not overlap as well, and the general lesson is worth stating: an architectural constraint is usually a statement about the systems layer of its year, not about the model.

  8. You must tune a model that activates one parameter in sixty-four, and you can afford only a dense proxy sweep. What do you do, and what do you tell your manager to expect?

    answer sketch

    Expect the transfer to be wrong, and say so up front. The September 2026 work on hyperparameter scaling across sparsity ran 1,800 pretraining runs and found that optimal learning rate and batch size both shift with the activation ratio in a way that neither the total nor the activated parameter count predicts. So a dense proxy gives you a starting point, not a setting. What I would actually do is sweep at the target sparsity on the smallest model I can afford, use the dense sweep only to bracket the range, budget for a short learning-rate re-search at scale, and watch the load histogram from step one, because a sparse model with a badly transferred learning rate does not fail loudly, it fails by collapsing the router.

What do you do with a list like this?

Not memorise the answers. The sketches are deliberately short because the point of each one is a single load-bearing idea, and if you have that idea you can build the sentences around it in the room.

The pattern across all three levels is the same. Level 1 is arithmetic you should be able to do on a whiteboard from the config alone. Level 2 is a question about which quantity is binding, and almost every one of them is answered by asking “at fixed what?” before answering. Level 3 has no settled answer, and the strong response names the axis it turns on and then commits to a default anyway.

That last habit is worth practising specifically. “It depends” is a bad answer on its own and a good one when it is followed by what it depends on and what you would do today.

Six parts ago the question was how a 671 billion parameter model trains for the price of a 37 billion parameter one. The answer turned out to be one edit to one sublayer, and five parts of consequences: a router with almost no gradient, a bank that collapses without supervision, a layer that spends more time on the network than on arithmetic, and a decode path that gives back most of its sparsity as soon as you batch.

None of that is hard. It is just longer than the one sentence the architecture is usually described in, and the gap between the sentence and the consequences is where the interview happens.