When the Plateau Is Still Learning: A Spectral Theory of Grokking
TL;DR for operators A model has already fit its training set, held-out performance is still poor, and the training curve has entered a long plateau. The usual operational question is whether more updates are buying anything. The paper examined here gives a more specific answer: some plateaued runs can still be changing in a way that prepares later generalization. Pracher and colleagues’ A Spectral Theory of Grokking: Weight Decay induces Feature Learning1 argues that coupled L2 weight decay leaves structured residual error after an initially near-fixed-feature fit. The remaining error is larger in task components that the model’s current geometry represents weakly. Under the paper’s theoretical conditions, that residual can then push the geometry itself toward those task-relevant directions. ...