Flow matching for language models had been sitting in my head as a buzzword, not a mechanism. I know diffusion. I know autoregression. But I could not explain, from first principles, why anyone would build a language model this way. So I decided to write it down properly: not a polished tutorial, but the notes I wish I already had, the kind I can come back to in six months and reload the whole picture in a single read.

That is the agenda for this series. Take flow based LLMs, strip them down to the ideas that actually matter, and lay them out from the continuous foundations up to the current state of the art. The mission is to simplify the topic for myself, and to leave a set of go-to notes I can revisit later when I need to reload the equations, the design choices, and the open problems.

This part is about one idea: generation is transport. Once that clicks, flow matching stops being a buzzword and becomes the obvious thing to do.

The bottleneck that motivates everything

An autoregressive language model emits one token per forward pass. To write 1000 tokens it runs the network 1000 times, each call strictly after the previous one. That is a hard sequential dependency. It is also why decoding latency, not training, is often the binding constraint when you deploy.

Now ask a different question. What if, instead of growing a sequence left to right, you started from pure noise (a random sequence) and gradually pushed the whole thing toward something that looks like data? Every position updates in parallel. The number of steps is a knob you control, not a function of the sequence length.

That is the diffusion idea, and it is also the flow matching idea. They turn out to be close relatives. To see why, I need to stop thinking about tokens for a moment and think about probability massProbability massOne unit of weight spread over outcomes, each non-negative, summing to one. For a discrete variable this is the probability mass function (PMF):\(p(x) = \Pr(X = x) \ge 0, \quad \sum_{x} p(x) = 1\)For continuous values the sum becomes an integral, $\int p(x)\,dx = 1$. Either way it is one fixed unit of weight, and flow matching transports it from noise onto the data..

Generation is moving mass from noise to data

Fix a simple source distribution $p_0$ that you can sample trivially, say a standard Gaussian. Fix the data distribution $p_1$ that you wish you could sample. Generative modeling is the problem of getting from one to the other.

Flow matching takes this literally. Imagine a continuous dial $t$ running from $0$ to $1$, and a family of distributions $p_t$ that smoothly morphs the source into the data:

\[p_t : \quad p_{t=0} = p_0 \ (\text{noise}), \qquad p_{t=1} = p_1 \ (\text{data})\]

This family is called a probability path. Picture a cloud of probability mass at $t=0$ spread out as a Gaussian blob, and at $t=1$ collected onto the data manifold. In between, it flows.

The word “flows” is doing real work. If mass moves continuously, it must move with some velocity. So picture, at every time $t$ and every point $x$ in space, a little arrow $u_t(x)$ that says which way the mass sitting at $x$ is headed, and how fast. That field of arrows, called the velocity field, is the only object we need.

Generating a sample is then almost embarrassingly simple: drop a noise point at $t=0$ and follow the arrows. Take a small step in the direction the local arrow points, look up the arrow at your new location, step again, and keep going until $t=1$. Wherever you land is a generated sample. In symbols, one such hop is an Euler stepEuler stepThe simplest way to follow a differential equation numerically: from your current point, move a small step of size $\Delta t$ along the local velocity, then repeat from where you land.\(x_{t+\Delta t} \approx x_t + \Delta t \cdot u_t(x_t)\)Smaller $\Delta t$ hugs the true path but costs more steps; larger $\Delta t$ is faster but cuts corners. This is forward Euler, the first ODE solver anyone learns.:

\[x_{t + \Delta t} \;\approx\; x_t + \Delta t \cdot u_t(x_t)\]

Shrink the step size toward zero and this follow-the-arrows recipe becomes an ordinary differential equation. The trajectory it traces out, as a function of the starting point, is called the flow $\psi_t$:

\[\frac{d}{dt}\,\psi_t(x) = u_t\big(\psi_t(x)\big), \qquad \psi_0(x) = x\]

Do not let the notation hide the picture. Both equations say the same thing: a noise point rides the field from $t=0$ to $t=1$, and where it ends up is your sample. No left to right token loop, and the number of steps is yours to choose.

Key idea A noise point rides the field from $t=0$ to $t=1$, and where it lands is your sample. The number of steps is yours to choose.
A Gaussian source morphing into a two-mode target along a probability path, with the velocity field that transports it. Original animation built for this post; the flow matching framework it illustrates is due to Lipman et al. (2023), Albergo and Vanden-Eijnden (2023), and Liu et al. (2023).

So if I just knew that field of arrows, I would be done. The whole problem collapses to one question: how do I get the arrows.

Learning the arrows, and the obstacle in the way

A velocity field is just a function. Hand it a point $x$ and a time $t$, it returns an arrow. We learn unknown functions all the time, and the workhorse is regression: collect input and target pairs, then fit a network to minimize the squared error between its output and the target.

So imagine someone handed me a training table. Each row is a point $X_t$ somewhere on the path at some time $t$, paired with the correct arrow $u_t(X_t)$ at that spot. Learning the field would be a five minute exercise: fit a network $u_t^\theta$ to that table by least squares.

\[\mathcal{L}_{\mathrm{FM}}(\theta) = \mathbb{E}_{t,\, X_t \sim p_t}\, \big\lVert u_t^\theta(X_t) - u_t(X_t) \big\rVert^2\]

This is the Flow Matching objective, and it says exactly what you would expect: sample a time and a point on the path, then nudge the network’s arrow toward the true arrow there. Nothing exotic. It is ordinary supervised regression onto a velocity.

There is one problem, and it is fatal as written. I do not have that training table. The true arrow $u_t(x)$, the marginal velocityMarginal velocityMarginal is the statistics word for averaging a variable out. A marginal distribution integrates over everything you are not asking about:\(p(x) = \int p(x \mid y)\, p(y)\, dy\)So the marginal velocity $u_t(x)$ is the conditional velocity $u_t(x \mid x_1)$, which aims at one specific data point $x_1$, averaged over all of them. The single arrow at $x$ blends the pulls of every plausible data point, and that averaging is exactly why it has no closed form., has no closed form. The reason is that it secretly averages over the whole dataset. The path $p_t$ is built by smearing a simple per-point path across every data point, and the velocity inherits that smearing, so the arrow at $x$ is a blend of the directions that every plausible data point would pull it. That blend runs over all of $p_1$, the distribution I am trying to learn and only ever see as samples. It is not a formula, it is an intractable integral. I can write the loss down, but I cannot evaluate the target it regresses onto. So close, and completely stuck.

The fix is the single most important idea in this whole literature, and it recurs in every paper I read. Instead of defining the path directly, build it from simple conditional paths, one per data point.

Pick a data sample $X_1 \sim p_1$. Now define a conditional path $p_{t \mid 1}(x \mid X_1)$ that interpolates from the source to that one specific data point. For a sensible choice, this conditional path and its conditional velocity $u_t(x \mid X_1)$ are available in closed form. The marginal path is then just the average over data:

\[p_t(x) = \mathbb{E}_{X_1 \sim p_1}\big[\, p_{t \mid 1}(x \mid X_1) \,\big]\]

The beautiful part is what happens to the velocity. The true marginal velocity turns out to be the posterior-weighted average of the conditional velocities:

\[u_t(x) = \mathbb{E}\big[\, u_t(X_t \mid X_1) \;\big\vert\; X_t = x \,\big]\]

I want to stress this equation because it is the load-bearing wall of the series. The thing I cannot compute, the marginal velocity, is a conditional expectation of things I can compute, the per-sample conditional velocities, weighted by the posterior $p(X_1 \mid X_t = x)$ of which data point a given noisy point came from.

Key idea The trainable velocity is always the posterior expectation of the conditional velocities. The thing you cannot compute is an average of things you can.

Stare at that for a second, because it is the whole trick. Fix a point $x_t$ in the middle of the path. Every data point proposes its own simple conditional arrow there, and the marginal arrow I want is their weighted average. I cannot afford that average, because it sums over the entire dataset. But an average does not have to be computed in one shot: it can be sampled. To make one training example I draw a single data point $x_1$, jump to a point $x_t$ on its closed-form conditional path, and read off that one conditional arrow. It is a one-sample estimate of the marginal arrow at $x_t$, usually wrong but right on average. And least squares, handed targets that are noisy but correct in expectation, drives the network straight to their mean. The mean of the conditional arrows is the marginal arrow. So I regress on the conditional target instead of the marginal one, and the averaging I could not afford happens on its own, one cheap sample at a time:

\[\mathcal{L}_{\mathrm{CFM}}(\theta) = \mathbb{E}_{t,\, X_1 \sim p_1,\, X_t \sim p_{t \mid 1}(\cdot \mid X_1)}\, \big\lVert u_t^\theta(X_t) - u_t(X_t \mid X_1) \big\rVert^2\]

This is Conditional Flow Matching, and the small miracle is that it has the same gradient with respect to $\theta$ as the intractable $\mathcal{L}_{\mathrm{FM}}$. The argument is short enough to sketch. Expand both squared norms. The $\lVert u_t^\theta \rVert^2$ term is identical in both. The cross terms are $-2 \langle u_t^\theta(x), u_t(x)\rangle$ versus $-2\langle u_t^\theta(x), u_t(x \mid X_1)\rangle$, and taking the expectation over the posterior of $X_1$ makes these equal precisely because $u_t(x) = \mathbb{E}[u_t(x \mid X_1) \mid X_t = x]$. The leftover terms do not depend on $\theta$, so the gradients match. We train against a target we can write down, and we converge to the velocity we actually wanted.

This is the trick that makes the whole approach “simulation free.” We never integrate the ODE during training. We sample a time, sample a data point, sample a point on its conditional path, and regress. No sampling loop in the training inner loop.

Rectified flow: the straightest possible path

The framework leaves the conditional path open. The simplest possible choice is a straight line, and it has a clean story of its own in the rectified flow work of Liu et al. (ICLR 2023).

Draw $X_0$ from the source and $X_1$ from the data. Connect them with a straight line and move along it at constant speed:

\[X_t = (1-t)\, X_0 + t\, X_1 \qquad \Rightarrow \qquad \dot{X}_t = X_1 - X_0\]

The conditional velocity is just the constant direction $X_1 - X_0$. Nothing could be simpler. Plug it into Conditional Flow Matching: regress a network onto that direction. The optimum is the conditional mean

\[v^\star(x, t) = \mathbb{E}\big[\, X_1 - X_0 \;\big\vert\; X_t = x \,\big]\]

which is exactly the marginal-velocity identity again, specialized to the linear path.

Why call it rectified, and why care about straightness. If the learned trajectories were perfectly straight lines, a single Euler step would carry a noise sample exactly to a data sample. One step, done. So straightness is worth chasing, and the paper makes this precise with a straightness measure: the straighter the flow, the fewer integration steps it needs.

But the first flow you train is not straight, and the reason is worth seeing clearly. You paired noise and data at random, then drew a straight line between each pair. Those lines cross all over the place. At a crossing point two different pairs want to head in two different directions, but the velocity field is a function: it can only return one arrow per point. So it returns the average of the two, and the trajectory that actually follows the field has to bend around the conflict. The interpolation lines cross, but the trajectories the model integrates never do.

Reflow fixes this by re-pairing. Take the flow you just trained and run it: each noise sample, pushed through the ODE, lands on one specific data sample. That hands you a new set of (noise, data) pairs, and these pairs are already untangled, because the ODE sends each start to exactly one end with no conflicts. Now throw away the original random pairing and train a fresh straight-line flow on the new pairs. Its interpolation lines cross far less, so it comes out straighter. Run it again to get even cleaner pairs and train again. Each pass is one reflow, and the flow straightens a little more every time.

The reason re-pairing can only help is a clean theorem, and it is worth decoding rather than skimming. Put the two pairings side by side. The old pair is $(X_0, X_1)$: a noise point $X_0$ and the data point $X_1$ it happened to be matched with at random. The new pair is $(Z_0, Z_1)$, and only one thing changes between them. You keep the very same noise point, so $Z_0 = X_0$. What you swap is its partner: $Z_1$ is the data point the trained flow actually delivers when you release that noise point and follow the ODE to $t=1$, and in general it is not the original $X_1$. Same starting noise, a new and better matched endpoint. Both pairings still draw their noise from the same source and their data from the same dataset; the only difference is who is matched with whom. With that in hand, for every convex cost function $c$,

\[\mathbb{E}\big[\, c(Z_1 - Z_0) \,\big] \;\le\; \mathbb{E}\big[\, c(X_1 - X_0) \,\big]\]

Read $c(X_1 - X_0)$ as the price of one trip from a noise point to its partner data point, and the expectation as the average price over the whole dataset. The inequality says reflow never raises that average price, and it says so for every convex way of measuring price at once: squared distance, plain distance, any convex $c$ you like. The intuition is the crossing once more. Two crossing trips form an X, and swapping their endpoints to uncross them shortens the total length. That shortening is just the triangle inequality, which is exactly the property convexity captures. Reflow performs that uncrossing for free, so the average trip can only get cheaper and the paths only straighter. Pushed to the limit, the coupling approaches optimal transport and sampling collapses toward a single step.

Key idea Reflow never raises the average cost of a noise-to-data trip, for every convex way of measuring it. Crossings are wasted detours, and uncrossing them can only help.

Hold onto this idea. Few-step sampling is exactly where the discrete language model work will fight its hardest battles later in the series.

The whole recipe as a training loop

Everything in Part 1 fits in about forty lines of PyTorch, and writing them out is the fastest way to feel how little flow matching actually asks of you. There are only two pieces to build: a network that predicts the velocity arrow, and a loss that regresses it onto the straight-line target $x_1 - x_0$.

import math
import torch
import torch.nn as nn

# The velocity field u_theta(x, t): given a point and a time, return an arrow.
class VelocityField(nn.Module):
    def __init__(self, dim=2, hidden=128):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(dim + 1, hidden), nn.SiLU(),
            nn.Linear(hidden, hidden), nn.SiLU(),
            nn.Linear(hidden, dim),
        )

    def forward(self, x, t):                 # x: (B, dim), t: (B, 1)
        return self.net(torch.cat([x, t], dim=-1))

# Conditional Flow Matching: regress onto one sampled arrow x1 - x0.
def cfm_loss(model, x1):
    x0 = torch.randn_like(x1)                # a noise sample from p0
    t = torch.rand(x1.size(0), 1)            # a random time in [0, 1]
    xt = (1 - t) * x0 + t * x1               # the point on the straight path
    target = x1 - x0                         # the conditional velocity (constant)
    pred = model(xt, t)                      # the network's arrow at (xt, t)
    return ((pred - target) ** 2).mean()     # least squares is the entire loss

# A toy target so this runs as is: eight Gaussian blobs in a ring.
def sample_data(batch):
    angles = torch.linspace(0, 2 * math.pi, 9)[:-1]
    centers = 4 * torch.stack([angles.cos(), angles.sin()], dim=1)
    idx = torch.randint(0, 8, (batch,))
    return centers[idx] + 0.15 * torch.randn(batch, 2)

# Train: sample data, take one CFM step, repeat.
model = VelocityField(dim=2)
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
for step in range(5000):
    loss = cfm_loss(model, sample_data(512))
    opt.zero_grad(); loss.backward(); opt.step()

# Sample: drop noise at t=0 and follow the arrows with Euler steps to t=1.
@torch.no_grad()
def sample(model, n, steps=50):
    x = torch.randn(n, 2)                    # start from pure noise
    dt = 1.0 / steps
    for i in range(steps):
        t = torch.full((n, 1), i * dt)
        x = x + dt * model(x, t)             # one Euler step along the field
    return x

samples = sample(model, 2000)                # these now land on the ring

Read cfm_loss again, because it is every idea from the last few sections squeezed into five lines. Sample a data point, sample noise, slide to a random time on the straight line between them, and push the network’s arrow toward $x_1 - x_0$. There is no ODE solve in the training loop, no integral over the dataset, no marginal velocity anywhere. The averaging that recovers the true marginal field happens on its own, one sampled arrow per step, exactly as the diagram showed. And sample is nothing but the Euler step from the top of the post, run fifty times.

Run it and the loss falls while the two thousand points drift out of the noise and settle onto the ring. Better still, you do not have to run anything: below is that same recipe, reimplemented in TensorFlow.js to train live in your browser. The demo holds a fixed set of noise points and maps them through the network every frame, so you watch the same points migrate from a blob onto the target as the network learns. It is the training itself you are watching, not a canned animation. Pick a target with the tabs, a ring, two moons, or a heart. The white outline marks where the points should land. Use Start, Pause, and Restart to drive it; training stops once the shape forms, and the live loss curve below tracks it the whole way.

Now notice what is missing from all of it: there is nothing here about images, nothing about text, nothing about any particular kind of data. The whole thing balances on one quiet assumption, the three operations on the interpolation line: $(1 - t)\, x_0$, $t\, x_1$, and their sum. Scale a point by a fraction, add two points. For vectors that is free. So if this is all it takes, why is there an entire research field about doing it for language? The answer starts with a sibling you have probably already met.

The sibling you already know: masked diffusion

If flow matching still feels exotic, its sibling will not. Diffusion tells the same story with one twist: start from a fully corrupted sample and walk it back to data, a few steps at a time.

For images, the corrupted sample is Gaussian noise, and undoing it is the continuous diffusion behind every modern image generator. For language the corruption is not noise, it is masking. Take a sentence, hide a growing fraction of its tokens behind a mask symbol until nothing readable is left, and train a model to put them back. Generation runs that process in reverse: begin from a fully masked sequence and unmask tokens a few at a time, every position in parallel, in as many or as few steps as you choose. These are the masked diffusion LLMs, the non-autoregressive relatives of GPT: MDLM, LLaDA, and SEDD, the line of work this series is really chasing.

Now look at what flow matching and masked diffusion share. Both throw away the strict left to right loop. Both begin from a content-free state, noise in one case and all masks in the other, then transport it toward data over a controllable handful of steps. That shared shape is not a coincidence. Flow matching is the continuous-transport telling of the story, and masked diffusion is the discrete corrupt-and-restore telling, and by the end of this series the two collapse into one framework written over different state spaces. The point of starting with flow matching is that the continuous version is where the ideas are cleanest. Once the velocity, the conditional trick, and few-step sampling are in hand, the masked diffusion LLMs stop looking like a separate technology and start looking like the discrete edition of the same thing.

Key idea Flow matching and masked diffusion are one framework, written over two different state spaces: continuous transport on one side, discrete corrupt-and-restore on the other.

There is just one obstacle between the clean continuous picture and the discrete models we actually want, and it is the obstacle masked diffusion was built around in the first place.

Why none of this works for text yet?

That obstacle is discreteness. Everything above lives in $\mathbb{R}^d$. The interpolation $(1-t) X_0 + t X_1$ assumes you can add a fraction of one point to a fraction of another and land somewhere meaningful. For an image, blending pixel values is fine. For a token, it is nonsense. There is no token halfway between “cat” and “dog.” The vocabulary is a discrete set, not a vector space, and the straight line between two one-hot vectors passes through points that are not valid tokens at all.

Key idea There is no token halfway between "cat" and "dog." The straight-line interpolation at the heart of flow matching has no meaning for discrete tokens.

That single sentence is the reason flow based language modeling is hard, and the reason the field exists as a separate research thread. The continuous machinery is elegant and well understood. Making it respect the discreteness of language is where the real design work happens.

There are two ways out. You can keep the continuous machinery and run it on the probability simplex, treating a token as a point in the space of distributions over the vocabulary. Or you can rebuild the whole framework natively in discrete space, replacing the velocity field with a jump process over tokens. Both routes lead to working language models, and they make very different tradeoffs.

That is where Part 2 picks up: the simplex relaxation, and why it both works and strains. For now the thing to carry forward is the one equation that will not change no matter how discrete things get: the trainable velocity is always the posterior expectation of conditional velocities. Hold that, and the rest is detail.