Cover image

The KL You Weren’t Watching: Moving RLVR Stability to the Query Side

TL;DR for operators RLVR teams face a familiar trade-off: stronger constraints can stabilize training, but constraining the response distribution too tightly can also suppress useful exploration. The paper identifies a second stability problem that response-side KL alone can miss. The same model being optimized also assigns probabilities to the training queries themselves, and those probabilities can shift substantially even when the dataset stays fixed. ...

September 28, 2026 · 8 min · Zelina
Cover image

Stale Rollouts, Fresh Trouble: The Two Speed Limits of Asynchronous RLHF

TL;DR for operators Asynchronous RLHF buys throughput by allowing rollout workers to continue generating completions while the learner updates the policy. The invoice arrives later: some rollouts were generated by a policy that the learner has already left behind. The paper’s useful contribution is not merely the familiar observation that stale data can destabilize training. It identifies two different speed limits.1 ...

July 19, 2026 · 20 min · Zelina
Cover image

Gated Sparse Attention: Speed Without the Sink

Context is expensive. That sentence is now obvious to anyone building with long-context models. The awkward part is that “long context” sounds like a capability, while the invoice often treats it as a lifestyle choice. Feed a model a 100-page contract, a repository, or a week of customer-support logs, and the theoretical promise is straightforward: the model can inspect more evidence before answering. The operational reality is less romantic. Attention cost grows quickly, prefill becomes painful, memory pressure rises, and training large models over long sequences can become unpleasantly dramatic. ...

January 24, 2026 · 17 min · Zelina