Cover image

Where You Pause Changes What You Forget

TL;DR for operators Kim and colleagues’ study of pause-token fine-tuning1 points to a training rule that is easy to miss if pause tokens are treated mainly as extra thinking time. At an equal pause-token budget, putting pauses at semantic boundaries is the only tested placement that consistently improves both math and code averages over ordinary supervised fine-tuning. The stronger variant also masks the loss on the pause tokens themselves. ...

October 1, 2026 · 6 min · Zelina
Cover image

Reward the Claim, Not Just the Answer: What V-Rubrics Changes in Multimodal RL

TL;DR for operators If a multimodal model reaches the correct final answer after misreading part of an image or making an invalid intermediate inference, a binary success reward gives the optimizer little information about what should actually be reinforced. The paper studied here turns those failure classes into separate training signals and, where possible, assigns their reinforcement-learning credit only to the relevant portion of the response. ...

September 21, 2026 · 7 min · Zelina
Cover image

Reward the Recheck: Reflection as a Control Surface for Reasoning Post-Training

TL;DR for operators Zhijie Wang’s GRPO and Reflection Reward for Mathematical Reasoning in Large Language Models1 offers two practical signals for teams tuning reasoning models. First, additional post-training stages do not compose automatically. On Qwen2.5-Math-7B, the base model scores 58.6% on MATH-500, supervised fine-tuning reaches 66.2%, and reinforcement learning applied directly to the base model reaches 73.8%. Yet the model that receives supervised fine-tuning before reinforcement learning falls to 57.6%. The authors attribute this weakness mainly to differences in the size and distribution of the datasets used at the two stages. ...

September 18, 2026 · 7 min · Zelina
Cover image

More Critics, Less Gain: Self-Questioning Has a Stability Limit

TL;DR for operators When a model can check its own reasoning, more self-checks are not automatically better. On GSM8K, Llama-3.2-1B rises from a 33.14% chain-of-thought baseline to 35.28% with one alternative critique and 35.84% with two, but falls back to 33.43% with three. The broader analysis links higher disagreement among these self-generated alternatives to greater reward variance and less stable policy updates. ...

September 15, 2026 · 6 min · Zelina
Cover image

The Reward Model Has to Move Too: R2M Tracks the Policy During RLHF

TL;DR for operators A reward model can remain technically unchanged while becoming operationally stale. As the policy shifts during RLHF, the evaluator increasingly scores outputs from a distribution different from the one that shaped its original preference fit. Real-Time Aligned Reward Model beyond Semantics proposes R2M, which lets the reward model incorporate current policy-internal representations and updates only a small fusion module plus the scoring head rather than retraining the full evaluator.1 ...

September 15, 2026 · 7 min · Zelina
Cover image

Shape the Signal, Keep the Objective: MeRLa’s Bet on Reusable RLHF Rewards

TL;DR for operators RLHF teams usually have two obvious levers: improve the reward model or improve the policy optimizer. Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback introduces a third.1 MeRLa learns an additional task-aware reward signal across auxiliary tasks, freezes it, and adds it to the existing reward during subsequent policy optimization. ...

August 23, 2026 · 7 min · Zelina
Cover image

Stale Rollouts, Fresh Trouble: The Two Speed Limits of Asynchronous RLHF

TL;DR for operators Asynchronous RLHF buys throughput by allowing rollout workers to continue generating completions while the learner updates the policy. The invoice arrives later: some rollouts were generated by a policy that the learner has already left behind. The paper’s useful contribution is not merely the familiar observation that stale data can destabilize training. It identifies two different speed limits.1 ...

July 19, 2026 · 20 min · Zelina
Cover image

The Reward Model Was Confident. That Was the Bug.

TL;DR for operators Reward models should not be treated as little oracles that hand down one clean number from the alignment heavens. In the paper’s diagnosis, the problem is more mundane and therefore more dangerous: a reward model can be wrong, uncertain, and numerically confident-looking at the same time. GRPO then standardizes those rewards inside a rollout group, giving extreme scores large influence even when the reward model is least reliable. Excellent. The pipeline has discovered a way to launder uncertainty into policy updates. ...

June 22, 2026 · 15 min · Zelina
Cover image

Rewarding Bad Physics Habits: What VLMs Learn When You Pay Them to Reason

A factory camera sees a pressure gauge. The AI reads the image, explains the mechanism, applies the formula, and recommends an action. Everyone in the meeting relaxes, because the model has produced a neat chain of reasoning. That is usually the moment to become nervous. The dangerous part is not that a vision-language model can be wrong. We know that. The more interesting problem is that a model can become wrong in a very specific way because we trained it to chase the wrong reward. Pay it for clean formatting, and it learns to look organized. Pay it for final answers, and it may sacrifice the reasoning path. Pay it to stare at the image, and it may do better on spatial problems while forgetting that physics also contains formulas. Apparently, “look harder” is not a complete theory of mechanics. ...

April 16, 2026 · 14 min · Zelina
Cover image

When Reasoning Pays (and When It Cheats): Fixing RL Signals in LLM Training

Scorecards are useful until people learn how the scorecard works. That is not a cynical observation. It is basic management. Sales teams optimize for commission rules. Customer-service teams optimize for handle-time dashboards. Students optimize for exams. And language models, with their charming lack of shame, optimize whatever reward function we put in front of them. ...

March 30, 2026 · 17 min · Zelina