Two distributions, and the chain between them

Fix notation. The world produces p. You collect D_0, a finite sample from p, and fit a model whose induced distribution is q_1. Sampling q_1 gives D_1; fitting that gives q_2, and so on. Each arrow loses something: fitting loses what the architecture cannot represent, sampling loses what a finite draw misses. Neither loss is recovered, because generation n+1 never sees p — only q_n’s output.

That separates the two regimes people conflate. A self-loop — a model training on its own outputs — has no external input, so errors only accumulate. A teacher pipeline, where a much stronger model or a symbolic checker generates the data, is a different object: the teacher holds information the student does not, so the arrow carries new bits in. Phi-style textbook corpora and verifier-filtered math and code data live in that second regime.

Advertisement

The recursion that collapses, worked

The cleanest derivation is one-dimensional Gaussian: generation n fits a mean and variance by maximum likelihood from M draws from generation n-1. The MLE variance divides by M, not M-1, so it is biased low by one degree of freedom — and that bias reapplies every round:

μ_n, σ_n = MLE on M draws from N(μ_(n-1), σ_(n-1)^2)

E[σ_n^2]  = σ_(n-1)^2 · (M-1)/M     ⇒   E[σ_n^2] = σ_0^2 · ((M-1)/M)^n

μ_n       = μ_(n-1) + σ_(n-1) · Z/√M,   Z ~ N(0,1)

Var(μ_n)  = σ_0^2 · [1 - ((M-1)/M)^n]  →  σ_0^2

The variance decays geometrically to zero — the distribution narrows to a spike — while the mean random-walks with accumulated variance converging to σ_0^2, so the spike lands about one original standard deviation from the truth. With M = 100 draws per generation the variance half-life solves ((M-1)/M)^n = 1/2, giving n ≈ 0.69·M ≈ 69:

(0.99)^100          = 0.366
σ_100 / σ_0        = √0.366 = 0.605      (60% of original spread)
Var(μ_100)          = σ_0^2 · (1 - 0.366) = 0.634 σ_0^2
RMS drift of mean   = √0.634 = 0.80 σ_0

Read the scaling, not the constants: the half-life is linear in M, so ten times more synthetic data per round buys ten times more rounds, not exponentially more. The danger therefore peaks exactly where synthetic data is most tempting — narrow domains where you generate a few hundred thousand examples and iterate.

Advertisement

Tail loss: what finite sampling deletes

The Gaussian model understates the problem because it assumes the family is correct. The sharper mechanism: rare events are precisely what a finite draw drops. An event of probability π is absent from n independent samples with probability (1-π)^n ≈ exp(-nπ).

π = 1e-6, n = 1e5   →  P(never sampled) = e^(-0.1) = 0.905
π = 1e-6, n = 1e7   →  P(never sampled) = e^(-10)  = 4.5e-5

Anything with π << 1/n vanishes from the next corpus, and a model fit where an event never occurs assigns it near-zero mass — so it is even less likely to survive the round after. Support shrinks monotonically, from the tail inward: rare dialects, unusual code idioms, minority facts, long-shot reasoning paths. Standard practice makes this worse on purpose, because generating below temperature 1 or with top-p truncation deletes the tail before you sample it — a corpus generated at T = 0.7, top-p = 0.9 arrives pre-collapsed. Mean validation loss will not show this; it is a tail statistic and must be measured as one.

Replace versus accumulate

The most consequential design choice is whether generation n trains on only the previous generation’s output or on the union of everything, real data included. The replace case is the recursion above, and it diverges. The accumulate case does not: the original real sample sits in the training set every round, anchoring the fit, so excess error stays bounded as generations grow instead of accumulating without limit.

There is an exact one-line reason. If the corpus is a mixture p_mix = (1-α)p + αq, then pointwise p_mix(x) ≥ (1-α)p(x) for every x, so:

KL(p || p_mix) = E_p[log(p/p_mix)] ≤ log(1/(1-α))

α = 0.5  →  ≤ 0.69 nats        α = 0.9  →  ≤ 2.30 nats

The bound holds no matter how degenerate q is. Real data does not merely dilute bad synthetic data; it puts a hard floor under the mass given to every region p cares about — exactly the tail protection the previous section demanded. Never delete the seed corpus.

Rejection sampling against a verifier

Filtering is not cleanup — it reshapes the distribution, which is why curated synthetic data can beat its own generator. Accept a sample with probability V(x) ∈ [0,1]; the accepted distribution is a reweighting:

q_acc(x) = q(x) · V(x) / a,      a = E_q[V] = Σ_x q(x)V(x)

cost: to keep D accepted tokens, generate N = D/a
a = 0.15, D = 2e9  →  N = 1.33e10 generated tokens

Two consequences. The acceptance rate a is the whole economics: filtering hard is a compute multiplier of 1/a on generation, so an aggressive filter can cost more than the training it feeds. And q_acc is a sharpened q, tilted toward whatever V rewards — if V is a correctness oracle it is supported only on correct outputs, so the student can exceed its own teacher on that axis. The coverage ceiling (rejection sampling can never produce what q assigns zero mass) is worked out in the verifier-scaling article and applies unchanged.

Verifiable domains behave nothing like open-ended ones

Everything above hinges on V, a completely different object in different domains. In math and code the verifier is near-perfect: run the unit tests, check the proof, compare against the known answer. Precision is essentially 1, the false-accept term is negligible, and q → q·V injects genuine external information — the environment’s verdict, which the generator did not have.

In open-ended domains — helpfulness, style, factuality — V is a learned reward model with precision ρ < 1. The accepted set carries a (1-ρ) fraction of bad examples, and those errors are not random: they are exactly the samples the reward model overrates. Train on them and the next generator produces more of that failure mode, which the same verifier again over-accepts. Errors correlate across rounds instead of averaging out — Goodhart’s law with a feedback loop. Rule of thumb: iterate where the verifier is an oracle; elsewhere treat filtered synthetic data as one-shot augmentation, not a loop.