TL;DR for operators

A model has already fit its training set, held-out performance is still poor, and the training curve has entered a long plateau. The usual operational question is whether more updates are buying anything.

The paper examined here gives a more specific answer: some plateaued runs can still be changing in a way that prepares later generalization. Pracher and colleagues’ A Spectral Theory of Grokking: Weight Decay induces Feature Learning1 argues that coupled L2 weight decay leaves structured residual error after an initially near-fixed-feature fit. The remaining error is larger in task components that the model’s current geometry represents weakly. Under the paper’s theoretical conditions, that residual can then push the geometry itself toward those task-relevant directions.

The result changes three training decisions. First, learning rate and weight decay should be interpreted jointly rather than tuned as independent controls. Second, extra post-fit updates can be meaningful when internal task-aligned structure is still evolving. Third, a targeted alignment intervention can be used diagnostically: in a paired 40-seed experiment, temporary early NTK alignment moved the mean first crossing of 50% held-out accuracy from 20,200 to 13,527.5 updates, with a paired mean speedup of 1.498.

That does not mean every plateau deserves more compute. Stronger decay can eventually make the relevant task mode unreachable, stronger still can prevent fitting, and large learning rates introduce a separate stability boundary. The paper provides a framework for distinguishing these regimes, not a rule to keep training indefinitely.

A fitted model can still retain structured supervision

Consider the point at which training accuracy has largely saturated. If the model were simply replaying a fixed representation, the remaining delay before generalization might look like slow optimization over already available features.

The paper starts from a different observation. With coupled L2 weight decay, an approximately fixed kernel does not drive every target component’s residual to zero. For an NTK eigendirection $j$, the equilibrium residual is

$$ \sigma_{j,\ast} = -\frac{D\lambda_W^{\mathrm{norm}}} {\Lambda_j + D\lambda_W^{\mathrm{norm}}} y_j. $$

Here, $\Lambda_j$ measures how strongly the model’s current local geometry represents that task direction. The smaller $\Lambda_j$ is, the larger the fraction of target signal left behind as residual.

This matters because the remaining error is not uniformly distributed noise. It is concentrated more heavily in task components that the current model geometry handles poorly.

To make that geometry concrete: around the model’s current parameters, some output changes are easy to produce with small parameter changes and others are difficult. The neural tangent kernel, or NTK, describes this local responsiveness. Decomposing it spectrally separates strongly represented directions from weak ones.

The paper’s key move is to treat the residual in those weak directions as continued supervision rather than as a mere artifact of imperfect fitting.

Residual error can move the geometry that produced it

A structured residual matters only if it can change the representation rather than merely decay under a fixed one.

Using neural tangent hierarchy dynamics, the paper derives a reduced feedback system for a selected task-aligned direction:

$$ \dot{\sigma} \simeq -(\Lambda+2\lambda_W^{\mathrm{norm}})\sigma -2\lambda_W^{\mathrm{norm}}y, $$
$$ \dot{\Lambda} \simeq -a\sigma -2\lambda_W^{\mathrm{norm}}\Lambda. $$

The first equation says that task-aligned tangent strength $\Lambda$ affects how quickly residual error relaxes. The second supplies the more consequential mechanism: with the required sign of the coupling coefficient $a$, the remaining residual can increase task-aligned tangent strength while weight decay contracts it.

Delayed generalization therefore becomes a competition between two processes. Residual supervision pushes the model toward a task-relevant geometry; decay limits how far that geometry can grow.

The modulo-97 MLP diagnostics support the existence of this post-fit movement. Fourier-like organization of the normalized label-space NTK continued strengthening after training accuracy had largely saturated. The internal state was still changing during a period that an accuracy-only dashboard could classify as a plateau.

Learning rate and decay jointly determine the slow clock

Once the fast residual dynamics are assumed to remain near their instantaneous equilibrium, the paper reduces the later stage to a slow coordinate

$$ \tau := \lambda_W^{\mathrm{norm}}t \simeq \eta\lambda_W s, $$

where $\eta$ is the learning rate, $\lambda_W$ is weight decay, and $s$ is the update count.

Away from critical boundaries, this implies a transition time proportional to

$$ (\eta\lambda_W)^{-1}. $$

The optimizer-plane experiments are substantial: 7,560 homogeneous MLP runs across an 84-by-90 learning-rate/decay grid, followed by 1,890 one-block Transformer runs across a 42-by-45 grid. The MLP sweep recovers the broad predicted phase organization and inverse-product timing dependence. The Transformer, despite violating exact homogeneity, shows a similar macroscopic memorization region, grokking band, failure regimes, and inverse-product timing pattern.

There is an important identification limit here. The $(\eta\lambda_W)^{-1}$ law alone does not establish feature learning. The paper explicitly notes that fixed-feature grokking accounts can generate similar timing dependence. What distinguishes its mechanism is the combination of the timing law with continued task-aligned NTK evolution and the intervention evidence.

Nor should weight decay be read as a monotonic accelerator. The reduced theory predicts several different boundaries: a finite-training boundary determining whether the task mode develops quickly enough within the update budget; a reachability cutoff beyond which decay prevents the task-aligned mode from reaching the generalization threshold at all; a stronger-decay fitting cutoff; and a separate large-learning-rate stability edge.

The intervention tests whether geometry is merely correlated

The paper’s temporary alignment experiment is especially informative because it changes the internal geometry and then removes the auxiliary objective.

Forty student runs were paired by initialization and data. For the first 3,000 updates, intervention students received an additional objective pushing their label-space NTKs toward a teacher’s NTK. After that point, both intervention and baseline students continued under the original training objective.

The mean first crossing of 50% held-out accuracy moved from 20,200 updates in the baseline condition to 13,527.5 under the intervention. The paired mean speedup was 1.498 and the median paired speedup 1.489.

This supports the direction of the paper’s mechanism: establishing task-aligned tangent structure earlier can advance later generalization even after the alignment pressure disappears.

It is not a clean test of the scalar one-mode model. The intervention changes several Fourier components simultaneously. Its role is therefore directional causal evidence for the relevance of early tangent geometry, not a proof that one spectral coordinate fully explains the transition.

For training teams, a plateau becomes a diagnosis problem

The operational lesson is narrower than “train longer.”

Paper evidence Operator interpretation Boundary
Weak NTK directions retain larger post-fit residuals Inspect whether remaining error is structured around poorly represented task components Derived under homogeneous squared-loss training with coupled decay
Task-aligned NTK structure can continue changing after training accuracy saturates Training accuracy alone may be insufficient to decide whether a plateaued run is finished Demonstrated on synthetic modular addition
Slow time scales approximately with $\eta\lambda_W s$ Treat learning rate, decay, and remaining update budget as a joint control problem Timing law alone does not identify feature learning
Stronger decay creates reachability and fitting cutoffs Increasing decay can move a run into a qualitatively different failure regime Boundary locations depend on reduced-model assumptions and calibration
Temporary NTK alignment advances later generalization Representation-alignment perturbations can test whether internal geometry is causally relevant Intervention affects several spectral components at once

For a training-system team, the inferred workflow is therefore to supplement loss and accuracy with measures of residual structure and task-aligned Jacobian or tangent geometry when deciding whether a stalled run deserves more compute. If those internal quantities are still moving in the relevant direction, additional updates have a mechanistic justification. If the task mode has stopped progressing, more elapsed training time is a weaker argument.

That recommendation remains a research proposition rather than a general production recipe. The exact derivations assume homogeneous networks, squared loss, coupled L2 decay, full-batch gradient flow, and an approximately decoupled task direction. Modular addition also provides unusually clean Fourier coordinates. The Transformer experiment shows that the macroscopic optimizer pattern can survive outside exact homogeneity; it does not establish that attention models follow the same microscopic residual-to-NTK mechanism.

The useful signal is what still changes after fitting

Grokking is visually striking because held-out accuracy can appear to change suddenly after a long period of successful memorization. The paper’s spectral account shifts attention from the jump itself to the quieter process before it: residual error remaining in weakly represented task directions and gradually changing the geometry through which the model learns.

For operators, that makes the post-fit plateau a state to measure rather than a duration to guess about. The relevant question is not simply how long the run has been flat. It is whether task-relevant internal structure is still moving, whether the optimizer settings allow that movement to reach the required threshold, and whether a controlled perturbation can demonstrate that the geometry actually matters.

Cognaptus: Automate the Present, Incubate the Future.


  1. Lenz Pracher and Pascal de Jong and Oskar Lieshaus and Alan Jeffares and Steffen Rulands (2026). A Spectral Theory of Grokking: Weight Decay induces Feature Learning. arXiv:2609.26679. https://arxiv.org/abs/2609.26679 ↩︎