TL;DR for operators
RLVR teams face a familiar trade-off: stronger constraints can stabilize training, but constraining the response distribution too tightly can also suppress useful exploration. The paper identifies a second stability problem that response-side KL alone can miss. The same model being optimized also assigns probabilities to the training queries themselves, and those probabilities can shift substantially even when the dataset stays fixed.
In the paper’s stability ablation, GRPO records a Policy-KL of only 0.0601 but a Query-KL of 0.9679. A run can therefore look restrained by its response-side metric while the model’s own likelihood distribution over the queries has moved much more.
Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization1 targets this second drift channel with Environment-Regularized Policy Optimization, or ERPO. It replaces response-side Policy-KL with Query-KL—the divergence between the current and reference model-induced query distributions—and adds reference-derived query weighting.
The mechanism matters because the Query-KL gradient contains the query log-likelihood gradient but not the response-policy score function. Explicit drift control can therefore move to the input side without directly spending the same response-side exploration budget, although shared parameters still create indirect coupling.
In the reported mathematical-reasoning experiments, ERPO raises mean Avg@32 from 0.274 to 0.336 and mean Pass@1 from 0.275 to 0.332 versus GRPO, while reducing the reported average train-evaluation gap from 6.47 to 3.14 percentage points. Long-horizon tests show smaller degradation than GRPO but not immunity to eventual collapse.
For RLVR teams, the practical implication is to monitor query-side drift separately from response-side drift and qualify checkpoints across decoding temperatures. The evidence remains narrow: mathematical reasoning, Qwen-family models, limited coefficient tuning, and no comprehensive repeated-seed uncertainty analysis.
A small Policy-KL does not mean the whole training process stayed close
KL regularization in RL post-training is usually understood as a restraint on the response policy. The trade-off is familiar: constrain the model too weakly and optimization can become unstable; constrain it too strongly and useful exploration can be suppressed.
The paper’s Table 3 exposes a problem with treating that response-side signal as a sufficient stability measure. Under GRPO, Policy-KL is 0.0601 while Query-KL is 0.9679. These numbers are not directly interchangeable measures of one phenomenon, but their coexistence establishes the key empirical point: relatively limited movement in the response distribution can occur alongside much larger movement in the model’s likelihood distribution over queries.
That query distribution needs careful interpretation. The training dataset itself is still fixed. The model is not suddenly sampling a different set of prompts. Instead, the same model being optimized also assigns autoregressive probabilities to the query sequences. Those probabilities change as the model parameters change.
The paper defines this model-induced environment shift as
where $\rho_\theta(q)$ is the current model’s induced distribution over query sequences and $\rho_{\theta_0}(q)$ is the corresponding distribution under the frozen pre-RL reference model.
This gives RLVR teams a second drift channel to inspect. Response-side KL tells you how the answer policy is moving. Query-side KL tells you how the optimized model’s representation of the training environment, expressed through query likelihood, is moving.
ERPO moves explicit regularization to the query side
ERPO makes two changes around an existing policy-gradient estimator.
The first is Query-KL. Instead of applying the explicit KL penalty to generated responses, the method penalizes divergence between the current and reference model-induced query distributions.
The second is reference-derived query reweighting. Queries that are more typical under the pre-RL reference receive different influence during optimization through a cached, dataset-static weighting scheme. The idealized importance weight is
The practical implementation uses a bounded proxy based on cached reference query scores rather than requiring exact density ratios.
The ablation results suggest these two components do different jobs. Query-KL provides the larger performance contribution. In the Qwen-7B, eight-rollout comparison, adding Query-KL to GRPO produces the strongest mean result over temperatures up to 1.0 among the reported ablations. Query weighting, by contrast, is associated mainly with lower Policy-KL and lower entropy, consistent with a stabilizing role.
That separation matters operationally. ERPO is not simply two regularizers added for the same purpose. One component primarily limits query drift; the other changes how much influence different queries exert on stochastic updates.
The gradient explains what ERPO does not directly constrain
A query-side penalty would be less interesting if it were merely another route to suppressing response exploration. The paper therefore makes a structural claim about the gradient.
For Query-KL,
The response-policy score function used by the policy-gradient estimator does not appear directly in this expression. The explicit regularization pressure operates through query log-likelihood instead.
This is the paper’s main mechanism for separating stability control from response exploration. It is also where overinterpretation is easy. Query-KL does not make the response policy independent of the regularizer. Query and response computations share model parameters, so changing those parameters can still affect both. The narrower claim is that Query-KL does not impose a direct response-score gradient of the kind used by response-side Policy-KL.
Implementation is comparatively light. Reference query likelihoods are computed in advance and cached, while current query likelihoods are reused from the policy-gradient forward pass. The authors therefore characterize ERPO as requiring no additional forward passes.
Accuracy gains are accompanied by stronger consistency tests
The main benchmark results cover six mathematical-reasoning datasets: AIME24, AIME25, AMC, MATH500, Minerva, and OlympiadBench. Evaluation spans sampling temperatures from 0.1 to 1.5 rather than relying on a single decoding configuration.
| Aggregate metric | GRPO | ERPO |
|---|---|---|
| Mean Avg@32 | 0.274 | 0.336 |
| Mean Pass@32 | 0.575 | 0.611 |
| Mean Pass@1 | 0.275 | 0.332 |
Those gains are the primary comparative performance evidence. More informative for training operations, however, are the diagnostics around them.
In the paper’s train-inference consistency analysis, GRPO has an average gap of 6.47 percentage points, compared with 3.14 for ERPO. The authors interpret the larger gap as evidence of greater overfitting or reward hacking. At the final reported checkpoint in the trajectory table, GRPO retains training accuracy of 76.7 while its evaluation accuracy falls to 58.4 at one test protocol and 66.2 at another. ERPO’s corresponding evaluation figures are 78.4 and 80.2 with training accuracy of 81.46.
The long-horizon experiment provides a different stress test. Training is extended to 1,000 steps. GRPO first shows marked high-temperature degradation after roughly 400 steps and later deteriorates more broadly. ERPO degrades less, but it does not avoid collapse indefinitely.
These experiments should be read as robustness and failure-mode diagnostics, not as separate claims that ERPO solves every instability mechanism.
Multi-temperature qualification may be as useful as the algorithm
One practical inference extends beyond whether a team adopts ERPO.
A checkpoint evaluated at one decoding temperature can hide instability that becomes visible elsewhere. In the Qwen-7B eight-rollout setting, both GRPO and ERPO perform reasonably at lower temperatures, yet both deteriorate sharply at temperature 1.5. Increasing ERPO’s rollout count from eight to sixteen substantially improves its high-temperature result, while the Qwen-32B ERPO run remains much stronger across the reported temperature range.
For release qualification, that makes decoding configuration part of the test surface rather than a downstream serving detail.
A useful monitoring set would therefore include response-side KL, query-side KL, entropy, train-evaluation divergence, and performance across the range of temperatures actually permitted in production. This is a Cognaptus inference from the paper’s diagnostic design, not a deployment protocol tested by the authors.
The same distinction applies to estimator portability. Supplementary experiments report positive gains when ERPO-style regularization is applied to DAPO and RLOO: mean accuracy below temperature 1.0 increases by 10.24 percentage points for DAPO and 2.28 points for RLOO. That supports compatibility beyond GRPO within the tested setting. It does not establish estimator-agnostic superiority across RLVR generally.
The evidence supports a new control point, not a universal recipe
The experiments are concentrated on mathematical reasoning using Qwen2.5-Math-7B and Qwen2.5-32B. Transfer to dialogue, code, instruction following, multilingual systems, or other model families remains untested in the supplied evidence.
The regularization coefficient is not exhaustively tuned. Query-likelihood estimates may also behave differently with other data-selection systems or larger-scale training pipelines. Most importantly, ERPO reduces long-horizon deterioration rather than eliminating it.
There is also limited evidence about run-to-run variability. The paper reports broad comparative testing—six benchmarks, ablations, temperature sweeps, model-scale comparisons, rollout-count variation, long-horizon trajectories, and additional RLVR estimators—but not a comprehensive repeated-seed uncertainty analysis.
That leaves a bounded conclusion. The paper provides evidence that response-side KL is not the only drift signal worth controlling in RLVR. Query-KL offers a technically distinct control point, and the reported experiments show that moving explicit regularization there can improve both accuracy and several stability diagnostics without directly placing the same gradient constraint on response exploration.
For post-training teams, the immediate question is therefore less whether Policy-KL is good or bad than whether it is measuring the part of training drift that actually needs to be controlled.
Cognaptus: Automate the Present, Incubate the Future.
-
Xianlei Zhou and Xiangdi Meng and Yu He and Tianyu Qi and Shuyan Guan and Xianli Zhang and Jian Zhang and Xin Li and Qika Lin and Jun Liu (2026). Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization. arXiv:2608.23311. https://arxiv.org/abs/2608.23311 ↩︎