Part 1 ended at a wall. Flow matching moves points along straight lines, and a straight line between two tokens runs through nonsense: there is no word halfway between “cat” and “dog.” The whole elegant machinery assumes you can add and scale your data, and tokens refuse.
There are two ways around that wall. This part is about the first one, and it starts with a reframing that sounds almost too cheap to work: stop treating a token as a symbol, and start treating it as a probability distribution.
A token from a vocabulary of $K$ words is usually written as a one-hot vector, all zeros except a single 1 at the word’s index. But a one-hot vector is also a probability distribution: the one that puts all of its mass on a single word. And the space of probability distributions over $K$ outcomes is not discrete at all. It is a smooth, continuous, convex region, and a continuous region is exactly the kind of place flow matching knows how to move around in.
A token is a corner, the simplex is the room
That region has a name. The probability simplexProbability simplexThe set of all probability distributions over $K$ outcomes: vectors that are non-negative and sum to one.\(\Delta^{K-1} = \{\, x \in \mathbb{R}^K : x_i \ge 0,\ \textstyle\sum_i x_i = 1 \,\}\)For $K = 3$ it is a triangle; the corners are the one-hot tokens and the interior holds the soft, uncertain ones. $\Delta^{K-1}$ is the set of vectors in $\mathbb{R}^K$ with non-negative entries that sum to one. For $K = 3$ it is a triangle. The three corners are the one-hot tokens, the center is the uniform distribution, and every point inside is a valid categorical distribution over the three words.
The reframing buys us a continuous home. A token is a corner. A noisy token is a point in the interior. Generation becomes a familiar thing: start from some easy distribution on the simplex, say the uniform one at the dead center, and flow toward a corner, the actual data token. We are back to transport, just on the simplex instead of $\mathbb{R}^d$.
The catch is that the simplex is not $\mathbb{R}^d$. It has a boundary, the edges and corners, and the naive straight-line interpolation from Part 1 either walks straight off it or hugs it badly. The two papers in this part are two answers to the same question: what is the right way to flow on the simplex?
Dirichlet flow matching: noise that lives on the simplex
Dirichlet Flow Matching (Stark et al., ICML 2024) starts from the most natural choice of noise on a simplex. The Dirichlet distribution is the canonical distribution over the simplex, and its concentration parameters $\alpha$ control how peaked it is. The conditional path toward a target vertex $e_i$, the data token, is a family of Dirichlets whose $i$-th concentration grows with time:
\[p_t(x \mid x_1 = e_i) = \mathrm{Dir}\big(x;\ \alpha = \mathbf{1} + t\, e_i\big)\]At $t = 0$ this is $\mathrm{Dir}(\mathbf{1})$, the uniform distribution over the whole simplex: pure noise. As $t$ grows, $\alpha_i = 1 + t$ climbs while the others stay at 1, so the distribution concentrates harder and harder around the corner $e_i$. At large $t$ it is a sharp spike on the data token. That is the noising schedule, and generation runs it in reverse: spread out at the start, concentrated on the answer by the end.
The conditional velocity that generates this path has a closed form. It points from the current point straight at the target vertex, scaled by a time- and position-dependent factor:
\[u_t(x \mid x_1 = e_i) = C(x_i, t)\,(e_i - x)\]The scalar $C(x_i, t)$ is where the Dirichlet machinery shows up: it comes from how the regularized incomplete Beta function, the cumulative distribution of the Dirichlet’s marginals, changes as the concentration $1 + t$ grows. It is not pretty, but it is a function you can evaluate, which is all flow matching needs.
Now the trick from Part 1 returns, in a particularly clean form. The marginal velocity is the posterior-weighted average of these conditional velocities, one per possible target vertex:
\[\hat v(x, t; \theta) = \sum_{i=1}^{K} u_t(x \mid x_1 = e_i)\, \hat p(x_1 = e_i \mid x; \theta)\]Look at what the network has to predict: $\hat p(x_1 = e_i \mid x)$, the probability that the clean token is vertex $i$ given the current noisy point. That is a classifier. The denoiser is not regressing a vector field at all. It is doing ordinary classification over the vocabulary, exactly what a language model already does. The loss is plain cross-entropy:
\[\mathcal{L}(\theta) = -\,\mathbb{E}\big[\log \hat p(x_1 \mid x; \theta)\big]\]This is the quiet payoff of the simplex view. The velocity field, the thing that looked so foreign in Part 1, gets assembled for you from a softmax classifier. You predict which token, and the closed-form conditional fields do the rest.
There is a bonus. Because everything is differentiable and lives on the simplex, classifier-free guidance carries over almost verbatim. A linear relationship between the learned flow and the score makes both classifier and classifier-free guidance exact rather than approximate, so steering generation toward a target property is a clean linear combination:
\[\hat v_{\mathrm{CFG}} = \gamma\, \hat v(x, t, y; \theta) + (1 - \gamma)\, \hat v(x, t, \varnothing; \theta)\]On promoter DNA design, where the vocabulary is four nucleotides, Dirichlet flow matching beats the discrete-diffusion baselines and even an autoregressive model, and a distilled one-step version stays competitive. For its target domain, it works.
The trouble with corners
But four nucleotides is not thirty thousand words, and the simplex hides a problem that only bites as you push it.
The Dirichlet conditional field, $C(x_i, t)(e_i - x)$, aims along the Euclidean straight direction $e_i - x$, a literal arrow from the current point to the corner. In the open interior of the simplex that is fine. But the boundary is where most of the action is, because a confident token is almost a corner, and near the boundary the geometry of the simplex stops looking flat. The natural distance between two distributions blows up as you approach an edge, so a Euclidean straight line is no longer the honest notion of “toward the corner.” Dirichlet flow matching keeps its own fields well-behaved there, but it is still steering with a Euclidean ruler on a space that is not Euclidean.
The deeper issue is that we have been measuring distance on the simplex with the wrong ruler. Euclidean distance treats the simplex as a flat triangle sitting in $\mathbb{R}^K$. But the simplex is a space of probability distributions, and the natural way to measure how far apart two distributions are is not Euclidean at all.
Fisher flow matching: the right way to measure distance
Fisher Flow Matching (Davis et al., NeurIPS 2024) fixes the ruler. The natural metric on a space of probability distributions is the Fisher-Rao metricFisher-Rao metricThe natural way to measure distance between probability distributions, the same object behind natural gradients and the Cramer-Rao bound in statistics. On the simplex it divides each squared step by the local probability:\(g(p)[u, v] = \sum_i \frac{u_i v_i}{p_i}\)Dividing by $p_i$ makes movement expensive where probability is small, which is exactly what respects the boundary of the simplex., the same object that underlies natural gradients in optimization. On the simplex it is:
\[g_{\mathrm{FR}}(p)[u, v] = \sum_{i} \frac{u_i v_i}{p_i}\]The division by $p_i$ is the whole story. It says movement is expensive where probability is small: nudging a token from probability $0.001$ to $0.002$ is a bigger deal than nudging $0.5$ to $0.501$, because it doubles the odds. This is the ruler that respects the boundary, and it is the very same $1/p_i$ that blew up on us a moment ago, now treated as geometry rather than a nuisance.
The elegant move is to make that awkward metric disappear with a change of coordinates. Map each distribution to the square root of its entries:
\[\varphi: p \mapsto s = \sqrt{p}\]Since $\sum_i (\sqrt{p_i})^2 = \sum_i p_i = 1$, the image of the simplex under this map lands on the positive orthant of the unit sphere, and the Fisher-Rao metric becomes the ordinary round metric on that sphere up to a constant factor of four. Scaling the map to $2\sqrt{p}$ makes it an exact isometry. Either way the boundary singularity is gone: the corners of the simplex become ordinary points on a smooth sphere, with no cliffs.
On a sphere, the analogue of a straight line is a great circle, a geodesicGeodesicThe shortest path between two points on a curved surface, and the curved-space stand-in for a straight line. On a sphere, geodesics are great circles, the paths an airplane follows on long-haul routes.. So Fisher flow matching runs the Part 1 construction with geodesics in place of straight lines. The conditional path is the constant-speed great circle from a noise sample to the data vertex,
\[x_t = \exp_{x_0}\!\big(t\, \log_{x_0}(x_1)\big)\]and the conditional velocity is the geodesic direction, with the same posterior-expectation trick underneath,
\[u_t(x_t \mid x_0, x_1) = \frac{\log_{x_t}(x_1)}{1 - t}\]where $\exp$ and $\log$ are the sphere’s exponential and logarithm maps. The network is trained with the Riemannian version of the conditional flow matching loss, regressing onto this geodesic velocity in the sphere’s tangent space.
Two more things fall out for free. Choosing the Fisher-Rao geometry is the same as doing natural-gradient descent, which is the right preconditioner for minimizing the forward KL divergence: the geometry you need for clean generation is the geometry you wanted for clean optimization anyway. And pairing noise with data through an optimal-transport coupling, a minibatch Sinkhorn solve, straightens the geodesics and lowers the variance of the training target. That is the simplex version of the reflow idea from Part 1, the same instinct that crossings are wasteful detours.
What it costs
So the simplex route works. It is elegant, it reuses the classifier you already have, it gives exact guidance, and with the right geometry it has no boundary pathologies. Why is the field not done?
Because the simplex is per token, and it scales with the vocabulary. For DNA, $K = 4$ and the simplex is a tetrahedron. For language, $K$ is thirty thousand or more, and every single position in the sequence now carries a thirty-thousand-dimensional simplex, or a point on a thirty-thousand-dimensional sphere. The geometry that was so clean in three dimensions is unwieldy in thirty thousand, and the methods feel it: Fisher flow’s own language-modeling numbers sit behind a plain autoregressive transformer on the standard one-billion-word benchmark, a perplexity of about 22.4 against the transformer’s 20.9, and the authors name large vocabularies and long sequences as the open problem.
There is a subtler cost too. Both methods spend their entire lives in a continuous relaxation. The token is never actually discrete during generation. It is a point drifting through the interior of a simplex, only becoming a real token at the very end when you read off the nearest corner. You are paying for a continuous ODE solve, in a thirty-thousand-dimensional curved space, to produce something that was discrete all along.
Where this strains, and what comes next
The simplex relaxation is the answer that keeps the most from Part 1. It holds onto the velocity field, the conditional trick, the geodesic version of straight-line transport, even classifier-free guidance. The price is geometry: a curved, high-dimensional, per-token space that gets harder to handle exactly as the vocabulary grows toward real language.
So the natural next question is the one Part 3 takes up. What if we stop relaxing? Instead of smuggling tokens into a continuous space and flowing through it, what if we build the flow directly in discrete space, where a token is a token and “moving” means jumping from one symbol to another? That means giving up the velocity field for something new, a rate of jumps, and it turns out to be exactly where flow matching and the masked diffusion LLMs from Part 1 finally meet.
That is the native discrete route. It is where the series stops relaxing and starts speaking the language of tokens directly.