TL;DR for operators
RL post-training can make a model better and longer at the same time. The usual response is to add a length penalty or bonus to the reward, but that changes the signal from which the optimizer decides which sampled responses to reinforce and which to suppress.
Weng and colleagues’ QGLAS method1 separates those roles. Quality determines the directional learning signal first. Length can then increase the reinforcement magnitude only for shorter responses that quality already favors.
The difference is substantial in the paper’s experiments. On Qwen3-4B, at roughly matched compression, QGLAS reduced response length by 32.4% while retaining 102.0% of the aggregate quality gain produced by quality-only RL. GR3 and GRLC achieved similar compression but retained 75.5% and 74.9%. On GLM-4.7-Flash, QGLAS retained 98.4% at 31.9% compression, versus 70.2% for GR3 and 68.3% for GRLC.
For teams running RL-based post-training, the practical question is therefore not merely how large a brevity penalty to apply. It is whether the efficiency objective is allowed to change the direction of the quality-derived learning signal at all.
Length control can interfere with the learning decision itself
Generating fewer tokens is attractive for straightforward reasons: inference consumes less compute, latency can fall, and high-volume serving becomes cheaper. But response length is not an independent quality variable. Additional tokens may contain repetition, or they may contain the explanation, qualification, or instruction-following detail that made the response valuable.
That makes direct reward modification consequential.
In group-relative RL, sampled responses receive advantages that determine their directional treatment during optimization: positive advantages push the policy toward a response, while negative advantages push away from it. If length is inserted into the reward before those advantages are constructed, the modification changes not only an individual response’s score but also the group statistics against which responses are compared.
The paper’s fixed-rollout diagnostic shows the resulting instability particularly clearly under dense rewards. With GR3 shaping, the macro-average advantage sign-reversal rate across three reward sources is 11.73% under dense feedback, versus 0.30% for averaged binarized variants of the same reward information.
This experiment is diagnostic rather than a full dense-RL-versus-RLVR training comparison. Its contribution is narrower: under graded open-ended rewards, small relative quality differences leave more room for a length adjustment to change whether the optimizer treats a response as desirable or undesirable.
For a model provider, that is a signal-integrity problem rather than merely a penalty-calibration problem.
QGLAS lets length change magnitude, not direction
QGLAS—Quality-Gated Length Advantage Shaping—moves the length intervention downstream.
The method first computes the advantage implied by quality alone. It then adds a bounded, non-negative conciseness bonus only when two conditions hold:
- the response already has a positive quality-derived advantage; and
- it is shorter than the mean length of the positive-advantage responses in its rollout group.
The core update is:
Here, $A_i^q$ is the quality-derived advantage, $h_i$ determines whether and how strongly a response receives a conciseness bonus, and $\lambda_g$ controls the group-level strength of that bonus.
Because the method only adds non-negative bonuses to already-positive responses, it enforces:
for every sampled response.
That property is the paper’s quality-polarity invariance principle. A secondary efficiency objective may change how strongly the optimizer favors a response, but it cannot turn a quality-favored response into a suppressed one or rescue a quality-disfavored response merely because it is short.
This also clarifies what QGLAS is not. It does not declare shorter responses inherently better. Negative-quality responses receive no brevity rescue. Nor does the method freeze the full ranking among positive responses: conciseness can still reorder responses whose quality signal already places them on the favored side.
Preserving signs is necessary, but the ablations show it is not sufficient
A tempting conclusion is that sign preservation alone solves the problem. The paper’s ablations argue otherwise.
At approximately matched compression on Qwen3-4B, full QGLAS records 102.0% quality-gain retention. Removing positive-only gating lowers retention to 74.6%; removing one-sided shaping produces 83.4%; removing advantage-level shaping gives 87.9%.
More revealing is the fixed-strength variant. It retains QGLAS’s polarity-preserving structural constraints yet reaches only 82.0% quality-gain retention, with weaker compression than full QGLAS.
The missing element is adaptive magnitude.
QGLAS increases the influence of conciseness when quality-favored responses are relatively close together and reduces it when their quality separation is large. In other words, the method does not treat a one-token saving as equally valuable in every rollout group. Length receives more room to influence optimization when the quality evidence provides less reason to distinguish among already-favored responses.
For implementation teams, this separates two design problems that are easy to conflate: protecting the direction of optimization and deciding how aggressively efficiency should influence the remaining degrees of freedom.
The main result is a quality-length trade-off, not simply shorter output
The matched-compression comparisons make the practical effect concrete.
| Policy model | Method | Compression | Quality-gain retention |
|---|---|---|---|
| Qwen3-4B | GR3 | 32.9% | 75.5% |
| Qwen3-4B | GRLC | 31.6% | 74.9% |
| Qwen3-4B | QGLAS | 32.4% | 102.0% |
| GLM-4.7-Flash | GR3 | 33.6% | 70.2% |
| GLM-4.7-Flash | GRLC | 31.5% | 68.3% |
| GLM-4.7-Flash | QGLAS | 31.9% | 98.4% |
Quality-gain retention is measured relative to the improvement produced by quality-only RL over the base model. A value above 100% therefore does not mean that every benchmark improved. It means the macro-average score slightly exceeded the quality-only RL reference after normalization.
The paper also varies the source of training feedback on Qwen3-4B. Across a learned reward model, rubric-based judging, and LLM-as-a-Judge feedback, QGLAS achieves 28.2-32.4% compression while retaining 99.4-102.0% of the corresponding quality-only RL gain without reward-specific hyperparameter retuning.
Three Qwen3-4B training seeds reproduce the matched-compression pattern. Direct pairwise comparisons, a length-controlled Arena-Hard aggregation, and an alternative judge are robustness checks supporting the comparative result rather than separate claims about the method.
For post-training teams, audit whether efficiency changes reinforcement polarity
The business implication is most concrete for model providers already running group-relative RL.
Cognaptus inference: before tuning the strength of a token-efficiency incentive, inspect whether introducing it at the reward level changes the sign of quality-derived advantages. If it does, the efficiency objective is no longer merely shaping concision; it is changing which behaviors the optimizer treats as desirable.
QGLAS offers one implementation pattern for avoiding that coupling. It modifies advantages rather than requiring a new reward model or changes to base-model weights, so its natural integration point is the post-training optimization policy.
The expected value is not “short answers” in isolation. It is lower generated-token volume while preserving more of the improvement already purchased through quality-focused RL training. That distinction matters for assistants, writing systems, dialogue products, and decision-support applications where response length affects serving economics but aggressive brevity can remove useful content.
The method inherits the quality objective’s mistakes
QGLAS protects fidelity to the quality-derived learning direction. It does not establish that the quality objective itself is correct.
If the reward model systematically prefers undesirable behavior, preserving its advantage signs preserves those mistakes as well. The method therefore assumes that quality-only RL provides a useful reference optimization.
The empirical scope is also finite: two policy-model families, three principal open-ended evaluation subsets, and three tested training reward sources. Alternative reward-source experiments are limited to Qwen3-4B, and GLM-4.7-Flash configurations were trained once because of their higher compute cost.
Finally, aggregate retention can conceal benchmark-specific movement. The paper explicitly notes that QGR is a macro-average measure, and in the length-controlled Arena-Hard analysis the bootstrap intervals for QGLAS and quality-only RL overlap.
Within those boundaries, the paper identifies a concrete design failure and a correspondingly precise remedy: when quality and efficiency compete during RL, shortening pressure does not have to participate in deciding whether a response deserves reinforcement. It can be restricted to deciding how much extra preference an already-good shorter response receives.
Cognaptus: Automate the Present, Incubate the Future.
-
Zijun Weng and Zhongan Bi and Xuanang Gao and Xiaohui Hu and Shuangyong Song and Yongxiang Li and Kaidong Yu and Xuanjing Huang (2026). Quality Determines Direction, Length Shapes Magnitude: Length Control for Open-Ended Reinforcement Learning. arXiv:2609.34718. https://arxiv.org/abs/2609.34718 ↩︎