Part 2 made tokens continuous by force. It lifted each token onto the probability simplex, ran flow matching through a curved, vocabulary-sized space, and read off the nearest corner at the end. It works, but it is a lot of machinery to model something that was discrete all along, and the bill comes due as the vocabulary grows.
This part takes the other road. Stop relaxing. Keep tokens discrete, and ask what “flow” even means when there are no in-between points to flow through.
The answer forces us to give up the one object the whole series has leaned on: the velocity. A velocity is a direction and a speed, and a direction only makes sense in a space where you can take a small step. Between “cat” and “dog” there is no small step. So in a discrete space, transport cannot be a smooth drift. It has to be a jump: at some moment, the token simply switches from one symbol to another. The right description is not a velocity but a rate, a probability per unit time of jumping.
Generation as a jump process
The object that replaces the velocity field is a Continuous-Time Markov ChainContinuous-Time Markov ChainA random process that sits in one discrete state and, at random moments, jumps to another. Unlike a discrete-time chain, jumps can happen at any real-valued time, governed by a rate for each possible transition.It is the discrete-space counterpart of an ODE or SDE: instead of drifting continuously, the state holds still and then hops. (CTMC). Picture a single token sitting on one of $K$ vocabulary symbols. At each instant it has, for every other symbol $j$, a rateTransition rate and rate matrixA rate $R_t(x, j)$ is the instantaneous probability per unit time of jumping from symbol $x$ to symbol $j$. Over a tiny interval $\Delta t$, the chance of that jump is about $R_t(x, j)\,\Delta t$.Stacking the rates for every pair of symbols gives the rate matrix $R_t$, the discrete stand-in for a velocity field. Its diagonal $R_t(x, x)$ is the negative of the total outgoing rate, so each row sums to zero, which is the mass left over for the token to stay put. $R_t(x, j)$: the instantaneous probability per unit time of jumping from its current symbol $x$ to $j$. Collect these into a rate matrix $R_t$. Most of the time the token sits still; occasionally, governed by the rates, it jumps.
Two equations carry over from Part 1, with jumps in place of drift. Sampling is still “follow the dynamics,” now an Euler step that either stays or jumps:
\[p_{t + \Delta t \mid t}(j \mid x_t) = \delta(x_t, j) + R_t(x_t, j)\,\Delta t\]With probability $R_t(x_t, j)\,\Delta t$ the token jumps to $j$; with the rest it stays put. And the continuity equation, the law that the dynamics must actually generate the path, becomes the master equationMaster equationThe discrete version of the continuity equation: it says how the probability of being in each state changes over time as mass jumps between states.\(\partial_t p_t = R_t^\top p_t\)Also called the Kolmogorov forward equation. A rate matrix generates a probability path exactly when this holds., the discrete Kolmogorov forward equation:
\[\partial_t p_t = R_t^\top p_t\]This is the discrete echo of $\partial_t p_t + \nabla \cdot (p_t u_t) = 0$ from Part 1. The velocity field generated a continuous flow; the rate matrix generates a jump process. Same role, different space.
The same recipe, one more time
Now watch the Part 1 machinery reassemble itself, identical in spirit.
You still build the path from simple conditional paths, one per data token. Generative Flows on Discrete State-Spaces (Campbell et al., ICML 2024), often called Multiflow, uses the two natural choices of discrete corruption. The masking path reveals the data token with probability $t$ and otherwise holds it in a special MASK symbol:
\[p_{t \mid 1}(x_t \mid x_1) = \mathrm{Cat}\big(t\,\delta(x_1, x_t) + (1 - t)\,\delta(M, x_t)\big)\]The uniform path reveals it with probability $t$ and is otherwise uniform over the $S$ symbols of the vocabulary:
\[p_{t \mid 1}(x_t \mid x_1) = \mathrm{Cat}\big(t\,\delta(x_1, x_t) + (1 - t)\,\tfrac{1}{S}\big)\]Either way, the conditional rate that generates the path has a closed form:
\[R^*_t(x_t, j \mid x_1) = \frac{\mathrm{ReLU}\big(\partial_t p_{t \mid 1}(j \mid x_1) - \partial_t p_{t \mid 1}(x_t \mid x_1)\big)}{S\,p_{t \mid 1}(x_t \mid x_1)}\]and the marginal generative rate is, once again, the posterior expectation of the conditional rates over the data:
\[R_t(x_t, j) = \mathbb{E}_{p_{1 \mid t}(x_1 \mid x_t)}\big[R^*_t(x_t, j \mid x_1)\big]\]That is the load-bearing wall from Part 1, rebuilt in discrete space. The marginal rate I cannot compute is the posterior-weighted average of conditional rates I can, weighted by the denoiser’s guess at the clean token. And the denoiser is the same clean-token classifier as in Part 2, trained by minimizing the cross-entropy, the negative log-likelihood of the clean token:
\[\mathcal{L} = -\,\mathbb{E}\big[\log p_{1 \mid t}^\theta(x_1 \mid x_t)\big]\]Predict the clean token, and the closed-form rates do the rest. The recipe has not changed at all. Only the space it runs in has.
There is one genuinely new dial, and it is a nice one. Campbell et al. show that the generating rate is not unique: you can add any rate that leaves the path’s marginals fixed, a detailed-balance rate, without changing what you are sampling.
\[R_t^\eta = R^*_t + \eta\, R_t^{\mathrm{DB}}, \qquad \eta \ge 0\]The parameter $\eta$ is an inference-time stochasticity knob. At $\eta = 0$ the sampler makes the fewest jumps it can. Turning $\eta$ up injects extra back-and-forth that lets the process revisit and correct earlier tokens, at the cost of more compute. It is a single number that trades determinism for self-correction, set at sampling time without retraining.
Discrete Flow Matching: the general path
Discrete Flow Matching (Gat et al., Meta, 2024) pushes the same idea to its general form and scales it up. Instead of committing to a single masking or uniform corruption, it builds the per-token conditional path as an arbitrary convex mixture of base corruptions, with a designable scheduler $\kappa$:
\[p_t(x^i \mid x_0, x_1) = \sum_{j} \kappa_t^{i, j}\, w^j(x^i \mid x_0, x_1), \qquad \sum_j \kappa_t^{i, j} = 1\]The full sequence path is a coordinate-factorized mixture of these per-token paths, averaged over couplings of source and data. The same discrete continuity equation must hold, now written with a discrete divergence, the net flux of probability out of a state minus the flux in:
\[\dot p_t(x) + \mathrm{div}_x(p_t u_t) = 0, \qquad \mathrm{div}_x(v) = \sum_{z} \big[v(z, x) - v(x, z)\big]\]And in the denoiser parameterization, where the network predicts the clean token, the generating velocity takes a form that should look eerily familiar:
\[u_t^i(x^i, z) = \frac{\dot\kappa_t}{1 - \kappa_t}\Big[p_{1 \mid t}(x^i \mid z) - \delta_z(x^i)\Big]\]Compare it to the continuous denoiser velocity from the start of the series: a scalar schedule factor times “what the clean data should be, minus where you are now.” The continuous and discrete worlds are running the same equation in different alphabets.
Discrete Flow Matching also brings corrector sampling, a way to combine the forward generating rate with a backward one to clean up errors during generation, and it scales: the paper trains a 1.7-billion-parameter model on a code and text mixture, reported as the first non-autoregressive model to produce non-trivial code.
Where the two threads meet
Now the payoff the series has been pointing at since Part 1.
Look again at the masking path: reveal the data token with probability $t$, otherwise hold it at MASK. That is exactly the forward process of a masked diffusion language model. The corruption that MDLM, LLaDA, and SEDD call “masking” is one specific choice of conditional path in the discrete flow matching framework, and the reverse CTMC that unmasks tokens is the generative rate we just built. Masked diffusion is not a cousin of flow matching after all. It is discrete flow matching, with the masking path and a particular schedule.
This is why Part 1 could promise that flow matching and the masked diffusion LLMs would turn out to be the same thing. In continuous space, diffusion was flow matching with noise turned on. In discrete space, masked diffusion is flow matching with the masking path turned on. The choice of corruption, mask versus uniform versus something learned, and the schedule $\kappa$ are the knobs. The CTMC, the posterior-expectation rate, and the clean-token classifier are the shared skeleton underneath.
What it buys, and what it still costs
The native discrete route fixes Part 2’s main complaint. There is no curved, vocabulary-sized continuous space to integrate through, no boundary geometry to manage, no relaxation tax. A token is a token, generation is a jump process, and the number of sampling steps is yours to choose, in parallel across all positions.
The honest accounting has the same shape as before, though. Discrete Flow Matching at 1.7 billion parameters narrows the gap to autoregressive models on code and text but does not close it. The absolute quality still trails a strong autoregressive model, and the best samples still want many CTMC steps. The promise is throughput, generation that escapes the one-token-per-forward-pass tax, and the open question is how few steps you can get away with.
That question, and a deeper one, is where Part 4 goes. We now have three things that look suspiciously alike: continuous flow matching, continuous diffusion, and discrete flow matching, with masked diffusion living inside the last one. Part 4 asks whether they are three frameworks at all, or one. There is a single abstraction, the generator of a Markov process, that holds every model in this series as a special case, and once you see it, the boundaries between flow and diffusion, continuous and discrete, stop being boundaries at all.