Part 4 dissolved the last of the distinctions. Flow matching, diffusion, discrete flow, masked diffusion: one framework, a generator with three terms, a path, and a Bregman loss. Once the paradigm question is settled, two questions are left, and they are the ones that decide whether any of this matters in practice. How fast can these models generate? And how far past plain text can they reach?

This final part is the honest state of both. The good news is real and the open problems are real, and a series that spent four parts admiring the theory owes you a clear look at where it actually stands.

How fast: a handful of steps

The whole case for leaving autoregression behind was throughput. An autoregressive model spends one forward pass per token: a thousand tokens, a thousand sequential passes. A discrete-flow or masked-diffusion model updates every position in parallel and takes as many steps as you let it, so its cost is the number of steps, not the length of the output. That only wins if the number of steps is small, and for a long time it was not. The best samples wanted hundreds or thousands of function evaluationsFunction evaluations (NFE)One evaluation of the denoising network. Sampling integrates the process over many small steps, and each step costs one network forward pass, so the number of function evaluations (NFE) is the real measure of sampling cost.Autoregression needs one NFE per token. A parallel diffusion sampler needs one per step, however long the sequence, so cutting the step count is the whole game., far more than the one-per-token of autoregression.

FS-DFM (Monsefi et al., ICLR 2026) is the result that changed the arithmetic. It is a token-space discrete flow matching model built around one observation. The standard sampler uses the instantaneous rate, the velocity at the current instant, and takes a tiny step; to take a big step honestly you would have to integrate that rate over the whole interval, and that integral has a closed form. Replace the instantaneous scale with the interval-integrated one:

\[\bar g_{t,h} = \frac{1}{h}\int_t^{t+h} \frac{\dot\kappa(\tau)}{1 - \kappa(\tau)}\, d\tau = \frac{1}{h}\ln\frac{1 - \kappa(t)}{1 - \kappa(t + h)}\]

and the same denoiser velocity from Part 3 becomes step-aware, accurate even when the step $h$ is large:

\[\bar u_t(x, z) = \bar g_{t,h}\big[p_{1 \mid t}(x \mid z) - \delta_z(x)\big]\]

Train it with a budget-aware objective, the usual path loss for fine steps and a self-consistency distillation for coarse ones, and the step count collapses. FS-DFM reaches generative-perplexity parity with a 1024-step baseline using 8 steps, up to 128 times fewer evaluations, at a similar model size.

Key idea The expensive part was the step count. Integrating the rate over a whole interval, instead of taking tiny instantaneous steps, lets a few-step sampler match a thousand-step one: FS-DFM hits parity at 8 steps, up to 128 times fewer evaluations.

This is the distillation thread from Part 4 cashing out. The diffusion duality’s consistency distillation, FS-DFM’s interval integration, the corrector tricks of discrete flow matching: all of them aim at the same target, producing in a handful of steps what used to take a thousand.

How fast: in production

The clean academic result has a noisy production counterpart. Mercury (Inception Labs, 2025) is a commercial masked-diffusion language model, a scaled descendant of the masked diffusion LLMs from Part 1, and it is fast in the way that shows up on a bill. On an H100 its code models run at over eleven hundred tokens per second, around ten times the throughput of speed-optimized autoregressive models of comparable quality. The parallel-generation promise, the thing this whole series has been circling, is running in production right now.

The honest footnote is that we do not fully know why, or how far it goes. Mercury is a technical report, not an open model. The serving stack, the exact masking schedule, the step count, and the custom kernels are proprietary, the report itself frames scaling diffusion language models as an open challenge, and it leaves the full speed-versus-quality trade-off uncharacterized. The throughput is real and measured. The recipe behind it is not on the table.

Key idea The throughput prize is real and shipping. Mercury runs at over a thousand tokens per second, roughly ten times speed-optimized autoregression, though the recipe behind the number stays proprietary.

How far: one generator, many modalities

The other axis is reach. The generator view from Part 4 does not care what the tokens represent. Text tokens, image tokens from a discrete codebook, both interleaved: a jump process over a discrete state space is a jump process whatever the symbols mean. So the same machinery that generates text can, in principle, generate images and understand them, in one model.

FUDOKI (Wang et al., 2025) is that idea built out. It is a single unified model for both multimodal understanding and generation, built entirely on discrete flow matching over a continuous-time Markov chain, with text and image tokens sharing one generator. At 1.5 billion parameters, initialized from an autoregressive backbone, it matches comparable autoregressive unified models on understanding and generation benchmarks.

The interesting part is a capability autoregression structurally cannot have. An autoregressive model emits tokens left to right and can never take one back: an early mistake is permanent. A flow over tokens revisits every position at every step, so it can revise a token it already wrote. The model that generates by jumping can also jump back, which is exactly what a unified understand-and-generate model wants when a later part of an image contradicts an earlier one.

Key idea A generator over discrete tokens does not care whether the tokens are words or image patches, so one model can understand and generate across modalities, and unlike autoregression it can revise tokens it already emitted.

The landscape, honestly

Zoom out from the individual papers and the picture is a field that has gone from curiosity to credible in about two years. A 2025 survey catalogs the large diffusion language models that now sit near the autoregressive frontier. LLaDA, an 8-billion-parameter masked diffusion model trained from scratch, reaches parity with LLaMA-3 of the same size; Dream, a 7-billion model adapted from an autoregressive checkpoint, reportedly edges past both. These are not toys. They are general-purpose language models that happen to generate by unmasking instead of by predicting the next token.

And yet the survey is also where the honesty lives, because it is candid about what is still broken, and the open problems are not small. Latency is the first: despite the throughput numbers, real deployments still tend to need a stack of tricks, caching, parallel decoding, distillation, to actually beat autoregression, and the theoretical parallelism does not always survive contact with a serving system. The second is likelihood. The non-sequential denoising makes the exact probability of a sequence intractable, and that matters, because the entire alignment toolkit, the RLHF and preference optimization that turned raw language models into assistants, is built on having that probability. The third is reasoning: chain-of-thought is sequential by nature, and a model that unmasks every position at once has no obvious place to put step-by-step thought. And underneath all of it sit long context, where the bidirectional attention is quadratic and length extrapolation is immature, and scaling laws, the thing that made autoregressive models predictable to build, which for diffusion language models are barely charted.

Key idea The open frontier is concrete: latency in real serving, intractable likelihood that blocks clean alignment, reasoning under parallel decoding, immature long context, and scaling laws that are barely charted.

Where this leaves us

Five parts ago this started with a single reframing: generation is transport, moving probability mass from noise to data. From that one idea the series built a velocity field, then a loss that learns it from single samples, then a way to do the same on the simplex, then a way to do it natively over discrete tokens, and finally the realization that all of it is one object, a generator, wearing different clothes.

The reason any of it matters is the one constraint autoregression cannot escape: it generates one token at a time. Flow-based and diffusion-based language models trade that strict sequence for a process that updates everything at once, and pay for it in sampling steps. The whole research program is the fight to make those steps few enough that the trade is worth it. Right now, in narrow domains and in production code models, it already is. In the general case it is close, and the open problems are clear enough to name.

I started this because I could not explain, from first principles, why anyone would build a language model this way. I think I can now. Whether it becomes the default way, or stays the fast alternative for the workloads that need throughput, is the part still being written.