Cover image

Shorter Without Flipping the Signal: QGLAS Reworks Length Control in Open-Ended RL

TL;DR for operators RL post-training can make a model better and longer at the same time. The usual response is to add a length penalty or bonus to the reward, but that changes the signal from which the optimizer decides which sampled responses to reinforce and which to suppress. Weng and colleagues’ QGLAS method1 separates those roles. Quality determines the directional learning signal first. Length can then increase the reinforcement magnitude only for shorter responses that quality already favors. ...

October 7, 2026 · 7 min · Zelina