Knowledge Gradient Ascent
Working notes on what I am currently learning: generative models, efficient ML systems, and the occasional fight with a GPU cluster. I write them to understand things properly, and leave them here in case they are useful to you.
-
SEP 2026entry 05 · 5 topics
How to Prepare for an ML Interview in 2026
A running syllabus for ML interviews: the questions that actually get asked, answered from first principles, one topic at a time, with the diagrams and animations I wanted while I was preparing.
-
MAY 2026entry 04 · 5 parts
Catching up on Flow based LLMs
A five-part build from first principles: why generation is transport, how it survives contact with discrete tokens, and how flow matching, diffusion, and masked diffusion turn out to be one framework.
-
22 MAR 2026entry 03
-
MAR 2026entry 02 · 3 parts
Distributed Training from Scratch
A three-part, from-scratch tour of training large models across many GPUs: the batch-size hierarchy, data and model sharding with DDP and FSDP, and what actually breaks at 256 GPUs.
-
1 OCT 2023entry 01
SSH Keys Demystified: A Researcher's Guide to Cluster Access
Picture this. You’ve just been granted access to a shiny HPC cluster. Four GPUs waiting for you. You open your terminal, type ssh cluster: