Cover image

The KL You Weren’t Watching: Moving RLVR Stability to the Query Side

TL;DR for operators RLVR teams face a familiar trade-off: stronger constraints can stabilize training, but constraining the response distribution too tightly can also suppress useful exploration. The paper identifies a second stability problem that response-side KL alone can miss. The same model being optimized also assigns probabilities to the training queries themselves, and those probabilities can shift substantially even when the dataset stays fixed. ...

September 28, 2026 · 8 min · Zelina