Three parts in, we have collected what look like four different machines. Part 1 built continuous flow matching: a velocity field that drifts noise into data, with diffusion showing up as its noisy sibling. Part 2 bent tokens onto a curved simplex. Part 3 built discrete flow matching: a rate matrix that jumps one token into another, with masked diffusion sitting inside it. Continuous and discrete, drift and jump, flow and diffusion: a pile of names, and a nagging sense that they keep reaching into the same toolbox.

This part is where the names collapse. It happens in three steps, each a tighter bridge than the last. First a change of view that makes discrete generation look like score matching. Then a literal equivalence that makes a discrete process the shadow of a continuous one. And finally a single abstraction that holds every model in this series as a special case, and makes the word “versus” in “flow versus diffusion” stop meaning anything.

The score view: SEDD

In Part 3, generation in discrete space was a rate: at each instant, a probability per unit time of jumping. There is another way to look at the very same reverse process, and it is the discrete echo of the single most important object in continuous diffusion, the score.

In continuous diffusion, the reverse process is driven by the score $\nabla \log p_t(x)$, the gradient of the log density, which points toward higher-probability regions. A discrete space has no gradient, but it has a clean analogue. For two tokens $x$ and $y$, the concrete scoreConcrete scoreThe discrete replacement for the gradient of the log density. Instead of a direction, it is a set of ratios between the probabilities of neighboring states:\(\frac{p_t(y)}{p_t(x)}\)It answers the same question the gradient does, which nearby state is more probable, but between discrete symbols rather than along a continuous direction. is the ratio of their probabilities under the noised distribution, $p_t(y)/p_t(x)$. SEDD (Lou et al., ICML 2024) is built on the observation that the reverse-time rate depends on the forward process only through these ratios:

\[\bar Q_t(y, x) = \frac{p_t(y)}{p_t(x)}\, Q_t(x, y), \qquad y \neq x\]

where $Q_t$ is the forward rate matrix. So to reverse a discrete corruption you do not need the rate directly. You need the concrete score, and SEDD learns it with a loss tailored to ratios, a score entropy (a Bregman divergence built from $-\log$), made tractable in a denoising form exactly as denoising score matching is in the continuous case.

The result was a milestone: SEDD was the first non-autoregressive model to match GPT-2 on its own zero-shot perplexity benchmarks, the result that made discrete diffusion language models credible.

For us the important thing is structural. The reverse CTMC from Part 3 and SEDD’s score-driven reverse process are the same process, parameterized two ways: once by a rate, once by a score. The rate and the concrete score determine each other, since $\bar Q$ is just the ratio times $Q$. The rate view and the score view are not two methods. They are two coordinates on one object.

Key idea The concrete score, the ratio p_t(y)/p_t(x), is the discrete analogue of the gradient of the log density. The reverse rate of Part 3 and SEDD's score are two parameterizations of the same reverse process.

The literal bridge: the diffusion duality

SEDD makes discrete generation rhyme with continuous diffusion. The Diffusion Duality (Sahoo et al., ICML 2025) makes them the same process.

Take the discrete corruption that is not masking: the uniform process, where each token’s forward distribution interpolates toward the uniform distribution over the $K$ symbols of the vocabulary:

\[q_t(\cdot \mid x; \alpha_t) = \mathrm{Cat}\big(\alpha_t\, x + (1 - \alpha_t)\,\tfrac{1}{K}\mathbf{1}\big)\]

Now run an ordinary Gaussian diffusion on the one-hot vectors, the kind from continuous diffusion, $w_t \sim \mathcal{N}(\tilde\alpha_t\, x,\ (1 - \tilde\alpha_t^2) I)$, and at each time take the argmax over the $K$ coordinates. The duality is that this argmax follows exactly a uniform-state discrete marginal:

\[z_t = \arg\max(w_t) \sim \mathrm{Cat}\big(\mathcal{T}(\tilde\alpha_t)\, x + (1 - \mathcal{T}(\tilde\alpha_t))\,\tfrac{1}{K}\mathbf{1}\big)\]

The operator $\mathcal{T}$ translates the Gaussian noise schedule into the discrete one. In words: uniform-state discrete diffusion is the argmax of a Gaussian diffusion. The discrete process is the shadow the continuous one casts when you collapse it to its most likely symbol.

This is not just pretty, and the paper turns it into three concrete wins. You can carry tools across the bridge: a curriculum that trains the discrete model through its Gaussian parent cuts gradient variance and roughly halves training time. The discrete view is provably the better one to optimize, because its evidence lower bound is tighter than the Gaussian parent’s, so you train in discrete space even though the equivalence runs through the continuous one. And the bridge enables a discrete consistency distillation, the discrete cousin of the few-step samplers that distillation buys in continuous diffusion, pulling sampling down toward a handful of steps.

Key idea Uniform-state discrete diffusion is the argmax of a Gaussian diffusion. Continuous and discrete are not rival paradigms; one is literally the shadow of the other.

The apex: generator matching

SEDD rhymes the discrete with the continuous. The duality equates two specific processes. Generator Matching (Holderrieth et al., 2024) names the abstraction that contains all of them.

Every model in this series is a Markov process: it carries a simple distribution to data by evolving a state over time, with the future depending only on the present. And every time-continuous Markov process is completely described by a single object, its infinitesimal generatorInfinitesimal generatorA linear operator that captures a Markov process by saying how the expected value of any function of the state changes per unit time, starting from the current state.It is the single object that contains a drift, a diffusion, and a jump rate as three separate terms, and it determines the whole process. $\mathcal{L}_t$. The generator is the instantaneous rule for how any measurement of the state changes:

\[\mathcal{L}_t f(x) = \lim_{h \to 0} \frac{\mathbb{E}\big[f(X_{t+h}) \mid X_t = x\big] - f(x)}{h}\]

for a test function $f$. It is the “what happens next, per unit time” operator, and it pins down the process entirely.

A generator generates a probability path exactly when one equation holds, the Kolmogorov forward equationKolmogorov forward equationThe universal conservation law for Markov processes: it says the expected value of any test function evolves according to the generator.\(\partial_t\, \mathbb{E}_{p_t}[f] = \mathbb{E}_{p_t}[\mathcal{L}_t f]\)It becomes the continuity equation for a drift, the Fokker-Planck equation for a diffusion, and the master equation for a jump process.:

\[\partial_t\, \mathbb{E}_{p_t}[f] = \mathbb{E}_{p_t}[\mathcal{L}_t f]\]

This single equation is every conservation law we have met. It is the continuity equation when the generator is a drift, the Fokker-Planck equation when it is a diffusion, the master equation when it is a jump. Same law, written for whatever generator you plug in.

And here is the structural punchline. On $\mathbb{R}^d$, any generator decomposes into exactly three pieces:

\[\mathcal{L}_t f(x) = \underbrace{\nabla f(x)^\top u_t(x)}_{\text{flow}} \; + \; \underbrace{\tfrac{1}{2}\,\sigma_t^2(x) : \nabla^2 f(x)}_{\text{diffusion}} \; + \; \underbrace{\int \big[f(y) - f(x)\big]\, Q_t(dy; x)}_{\text{jump}}\]

Flow, diffusion, jump. That is the entire design space. Flow matching is the generator with only the first term. Diffusion keeps the second. Discrete flow matching and the masked diffusion LLMs are the third. They were never different frameworks. They were different terms of the same operator.

Key idea Every iterative generative model is a Markov process, and on continuous space its generator splits into exactly three terms: flow, diffusion, jump. Each method in this series is a choice of which terms to use.

Two more pieces complete the picture, and both are old friends.

The load-bearing wall holds in full generality. The marginal generator, the one you cannot compute, is the posterior expectation of conditional generators you can:

\[\mathcal{L}_t f(x) = \mathbb{E}_{z \sim p_{1 \mid t}(\cdot \mid x)}\big[\mathcal{L}_t^z f(x)\big]\]

This is the identity from Part 1, now stated for any generator at all. And the loss that makes it trainable is, once again, not arbitrary: the Bregman divergences are exactly the family for which regressing on the conditional generator has the same gradient as regressing on the marginal one. The squared-error loss of flow matching and the cross-entropy of the discrete models are both Bregman divergences. The “conditional matching” trick that has shown up in every part is one theorem, instantiated over and over.

Finally, generators add. A convex combination of two generators that solve the same forward equation is again a valid generator for the same path:

\[\mathcal{L}_t^{\mathrm{super}} = \alpha_t^{1}\, \mathcal{L}_t + \alpha_t^{2}\, \mathcal{L}_t', \qquad \alpha_t^1 + \alpha_t^2 = 1\]

so you can superpose flow and jump, run a model that drifts and hops at once, and it stays a principled generative model.

Key idea The conditional-matching trick is one theorem. For any generator, the marginal is the posterior expectation of conditionals, and any Bregman divergence trains it; flow matching's squared error and the discrete models' cross-entropy are two instances.

Redrawing the map

Step back, and the series looks different. There are not four methods. There is one recipe with three knobs. Choose a generator, which is choosing how much of flow, diffusion, and jump to use. Choose a probability path from noise to data, built as always from simple conditional paths. Choose a Bregman loss to regress the marginal generator onto the conditional one. Everything we have seen is a setting of those three knobs:

Method Generator State space
Flow matching flow (ODE) continuous
Diffusion flow + diffusion (SDE) continuous
Dirichlet / Fisher flow flow on simplex / sphere continuous (relaxed)
Discrete flow matching jump (CTMC) discrete
Masked diffusion LLMs jump, masking path discrete

Flow versus diffusion was never a rivalry. Continuous versus discrete was never a wall. They are coordinates in the space of generators, and the choice between them is an engineering decision about your data and your compute, not a clash of paradigms.

What is left

If it is all one framework, the interesting questions are no longer about which paradigm. They are how fast, and how far.

How fast: every model here generates by simulating a process, and simulation costs steps. The whole reason to leave autoregression behind was throughput, and throughput is steps times cost per step. Part 5 is about collapsing the step count, the few-step and one-step discrete samplers, the distillation tricks the duality unlocked, and the systems already running this in production.

How far: the generator view does not care whether the state is text, image patches, or both, so the same machinery reaches naturally toward any-to-any multimodal models. The frontier is models that read and generate across modalities with a single generator, and that is where the series ends.