Cover image

Poison Does Not Scale Once: Why Pre-Training Contamination Changes Regime

TL;DR for operators A pre-training contamination test can produce a clean-looking power law without producing a reusable contamination rule. In controlled OLMo-style runs, increasing the poison fraction approximately follows a power law in relative clean-validation perplexity degradation, but the fitted relationship changes over training and with model size. Halder, Dey, and Pehlevan’s A Solvable Theory of Pre-training Data Poisoning: Regime-Dependent Scaling Exponents1 explains why such movement may be structural rather than measurement noise. ...

October 8, 2026 · 7 min · Zelina