How to Prepare for an ML Interview in 2026
I am interviewing at ML companies this year, and the preparation splits cleanly into topics. Each topic here is a series I wrote while working through it: the questions an interviewer opens with, the derivation underneath, and the figures that make the answer stick. Start anywhere. More topics land as I get to them.
-
SEP 2026topic 01 · 6 parts
Interrogating the Transformer
The questions ML interviews actually ask about transformers, answered from first principles: one pass through the block, why the square root is there, what residuals and norms buy you, how position gets in, what the FFN knows, and what changed since 2017.
-
SEP 2026topic 02 · 6 parts
Interrogating Mixture of Experts
What a Mixture of Experts actually is and why every frontier model became one: total versus active parameters counted on DeepSeek-V3, how a token picks its experts, why routers collapse and how three different mechanisms stop them, what an MoE layer costs in memory and network traffic, and what the field still has not settled in September 2026.
-
SEP 2026topic 03 · 6 parts
Interrogating the KV Cache
The one object that decides how an LLM serves: why the cache exists, exactly how many bytes it is, every way people make it smaller, how vLLM and SGLang manage it, and what it does to tokens per second.
-
SEP 2026topic 04 · 7 parts
Interrogating the GPU
How a model actually meets the hardware: the four bottlenecks, the roofline, the memory hierarchy inside a GPU, where every byte goes during training and serving, FlashAttention from first principles, the September 2026 toolkit, and what to watch when building at scale.
-
SEP 2026topic 05 · 7 parts
Interrogating Parallelism
The five ways a model gets split across a cluster, derived one at a time: the two budgets that force it, data parallelism and the ZeRO ladder, tensor and sequence parallelism, pipeline schedules and their bubbles, context and expert parallelism, how the axes compose onto a real topology, and a graded question set.