TL;DR for operators

A pre-training contamination test can produce a clean-looking power law without producing a reusable contamination rule.

In controlled OLMo-style runs, increasing the poison fraction approximately follows a power law in relative clean-validation perplexity degradation, but the fitted relationship changes over training and with model size. Halder, Dey, and Pehlevan’s A Solvable Theory of Pre-training Data Poisoning: Regime-Dependent Scaling Exponents1 explains why such movement may be structural rather than measurement noise.

The paper’s strongest result is theoretical. In a solvable heavy-tailed model, the optimized degradation exponent changes depending on the data-to-dimension regime and on the order in which the ample-data and small-poison limits are taken. There is therefore no theoretical basis here for assuming one contamination exponent applies across regimes.

For pre-training governance, that changes the design of stress tests. A poison fraction measured at one model scale or checkpoint should not automatically be translated into expected degradation elsewhere. Comparisons need matched architectures, training horizons, optimization conditions, and evaluation streams.

The proposed transformer mechanism is more tentative. Finite training time may act as a cutoff that leaves weak-curvature directions partly learned and disproportionately exposed to contamination. The paper measures power-law-like behavior near the hard edge of a preconditioned Gauss-Newton spectrum, but this is supporting evidence for a mechanism, not a validated contamination sensor.

One poison fraction does not imply one degradation curve

The operational quantity is straightforward: compare a poisoned model with a matched clean model on clean evaluation data,

$$ \Delta(\epsilon)=\frac{P_{\mathrm{poisoned}}(\epsilon)}{P_{\mathrm{clean}}}-1, $$

where $\epsilon$ is the pre-training poison rate.

Across the controlled OLMo-style experiments, this degradation is approximately described over the observed range by

$$ \Delta(\epsilon)\approx C\epsilon^a. $$

The tempting move is to estimate $a$ once and turn it into a contamination rule: double the poison rate, infer the resulting quality loss, then reuse that relationship for another checkpoint or larger model.

The experiments do not support that simplification. The fitted slope and intercept drift with training progress and model size. The model configurations reported in the paper span roughly 135 million to 1.287 billion parameters, with training horizons ranging from 8.7 billion to 39.9 billion tokens. Clean and poisoned runs are matched on architecture, token budget, optimization schedule, and held-out clean evaluation.

That makes the moving exponent the central puzzle rather than a nuisance parameter.

Smooth contamination should have produced a quadratic response

The paper first establishes what cannot easily explain fractional scaling.

Suppose contamination changes the training objective smoothly around a well-behaved clean optimum. Under analyticity and nondegeneracy assumptions, the learned parameters move by order $O(\epsilon)$. Because clean evaluation risk is stationary at the clean optimum, the leading risk increase is then second order:

$$ \mathcal{R}(\theta_\epsilon)-\mathcal{R}(\theta_0) = \frac{\epsilon^2}{2} g^\top H^{-1}GH^{-1}g +O(\epsilon^3). $$

The generic exponent is therefore 2.

This result is useful because it narrows the search. A persistent non-integer exponent cannot simply be attributed to a small smooth perturbation of an otherwise regular optimum. Some singular or non-analytic structure has to enter the effective learning problem.

The authors introduce one such structure through heavy tails combined with an adaptive cutoff. If excluding extreme observations creates clean bias of order $\tau^{-\beta}$ while admitted poison influence grows as $\epsilon\tau^\gamma$, optimizing the cutoff gives

$$ a=\frac{2\beta}{\beta+\gamma}. $$

Heavy tails alone are not the claim. The fractional law comes from their interaction with a cutoff whose effective scale changes with contamination.

The exponent changes when the data regime changes

The paper then makes the mechanism exact in a high-dimensional, coordinate-separable ridge-regression model with Pareto-tailed clean covariates, label-shift poisoning, and coordinatewise truncation.

Its main result is that two asymptotic regimes produce different optimized exponents.

Regime Optimized clean degradation Mechanism
Fixed positive aspect ratio $\phi$ $\epsilon^{q_\star/(q_\star+2)}$ Poison can arrive as a rare hit on a coordinate with limited clean support, creating an $O(\tau)$ displacement and $O(\tau^2)$ squared error
Ample-data limit taken first $\epsilon^{2-2/q_\star}$ Many clean observations dilute the poison contribution into an $\epsilon\tau$ population shift before risk is squared

Here $q_\star$ is the Pareto tail exponent.

The distinction is not two alternative curve fits to the same situation. It comes from non-commuting limits. Keeping the high-dimensional aspect ratio positive while poison becomes rare leaves finite clean coverage per direction. Taking the ample-data limit first changes how poison is averaged within those directions.

Numerical simulations reproduce the predicted exponent branches for $q_\star$ values from 1.5 to 4 and show a crossover as the aspect ratio changes. Those simulations are checks of the theoretical construction rather than independent evidence that full transformers obey the same mechanism.

For operational readers, the significant point is regime dependence: contamination sensitivity can depend on effective data coverage relative to model directions, not only on the global percentage of poisoned tokens.

Finite training time is the proposed bridge to LLMs

The transformer connection begins with a more familiar training observation: at a finite checkpoint, not every parameter direction has been learned equally.

In the paper’s local quadratic approximation, with a frozen Adam preconditioner, training time $T$ produces the spectral response

$$ q_T(\lambda)=\frac{1-e^{-\lambda T}}{\lambda}. $$

Directions with larger curvature are learned more completely. Weak-curvature directions remain partially unresolved. Training time therefore behaves like a moving spectral cutoff.

Under an additional assumption that inverse curvature has sufficiently heavy-tailed structure, the same basic balance reappears: stopping earlier leaves clean residual bias, while continuing training allows more poison response to accumulate.

The authors measure the hard edge of the preconditioned Gauss-Newton spectrum in a 167.7-million-parameter OLMo-style model trained to 3 billion tokens. Within the numerically resolved window, they estimate a hard-edge exponent of roughly 0.30–0.33.

That result supports the plausibility of weak-curvature heavy-tail structure. It does not establish that this structure causes the poison-rate scaling seen in the pre-training experiments.

Contamination stress tests need matched regimes

What the paper directly shows: in its solvable model, contamination exponents are regime-dependent, and smooth regular perturbations alone generically imply quadratic rather than fractional degradation. The controlled LLM experiments also show that an approximate empirical power law is not stable across training time and scale.

Cognaptus inference for practice: organizations evaluating pre-training supply-chain risk should avoid treating a poison percentage as carrying a fixed quality penalty. Stress tests are more defensible when poisoned and clean models are compared at matched checkpoints, model scales, token budgets, and optimization schedules. A sensitivity estimate obtained halfway through training should not automatically become a tolerance rule for a later checkpoint or a different parameter scale.

This also argues for separating aggregate clean-quality degradation from targeted backdoor testing. The paper studies the former. A pipeline can therefore pass a trigger-oriented security check while still requiring a separate assessment of broad clean-model quality loss.

Curvature diagnostics remain a research hypothesis

The strongest evidence in the paper belongs to the stylized theoretical model. The transfer to full LLM pre-training depends on several additional assumptions: local quadratic dynamics, a frozen-preconditioner approximation, heavy-tailed inverse-curvature structure, and an appropriate relationship between training horizon and the relevant crossover scale.

The very-small-poison regime in real LLM training is also underexplored. The paper explicitly leaves open whether the observed empirical scaling changes once the poison rate approaches or falls below the effective aspect-ratio scale.

Curvature measurements therefore should not yet be treated as production contamination indicators. They are better understood as a candidate diagnostic for testing whether weakly learned directions help explain why apparently similar poison fractions behave differently across checkpoints.

The durable result is narrower and more consequential: contamination sensitivity need not belong to the dataset alone. It can depend on the training regime in which that contamination is encountered. Data-governance thresholds built from one experiment should preserve that context before being extrapolated elsewhere.

Cognaptus: Automate the Present, Incubate the Future.


  1. Indranil Halder and Rastri Dey and Cengiz Pehlevan (2026). A Solvable Theory of Pre-training Data Poisoning: Regime-Dependent Scaling Exponents. arXiv:2609.32288. https://arxiv.org/abs/2609.32288 ↩︎