Cover image

Reward the Right Thing: GUI Agents Need Better Success Criteria, Not Just Better Judges

TL;DR for operators After a GUI agent sends a message, edits a document, moves a file, changes a setting, or searches for information, someone—or something—has to decide whether the instruction was actually completed correctly. It is tempting to treat that decision as mainly a model-capability problem: use a stronger vision-language model, show it more screenshots, or improve the prompt. The evidence in Task-Adaptive Rubrics for GUI Reward Modeling suggests another failure point comes earlier. The verifier first needs an adequately specified definition of success.1 ...

September 29, 2026 · 7 min · Zelina
Cover image

When Preference Strength Becomes Part of the Model

TL;DR for operators Human reviewers often provide more information than “A is better than B.” They may say one answer is slightly better, another clearly better, and another much better. The operational question is whether the reward model should learn those distinctions directly or whether teams should translate them into hand-set margins, scales, or soft targets. ...

September 16, 2026 · 7 min · Zelina
Cover image

When the Wrong Label Shouts Loudest: Correcting Preference Noise in RLHF and DPO

TL;DR for operators Preference pipelines have an awkward failure mode: when an annotator accidentally chooses the worse of two responses, a standard preference loss can push hardest on exactly the pair the model ranks most strongly against the recorded label. A bad label can therefore receive unusually strong corrective force instead of being naturally ignored. ...

September 16, 2026 · 8 min · Zelina
Cover image

Your Reward Model Is Also a Voting Rule: What Social Choice Changes About RLHF

TL;DR for operators When thousands of annotators disagree about two acceptable model responses, the pipeline still has to decide how that disagreement becomes model behavior. That decision is often hidden inside familiar technical choices: how comparisons are sampled, whether votes are collapsed into majority labels, which reward-model class is fitted, and how aggressively a policy is optimized against the resulting score. ...

September 16, 2026 · 8 min · Zelina
Cover image

No Click Is Not a No: Turning Passive Feedback Into Reward Signals

TL;DR for operators If your product already has thousands of clicks, copies, likes, shares, and non-actions, the tempting shortcut is to use those events as cheap preference labels. The problem is that the same no-action can mean a disliked response or a satisfactory response from a user who simply did not act. The event log records behavior, not preference directly. ...

September 15, 2026 · 8 min · Zelina
Cover image

One Score, More Signals: Making Reward Models Easier to Rank and Audit

TL;DR for operators A system generating several candidate answers eventually needs a ranking decision: which response should be shown, which should be discarded, and which should receive additional review. A reward model commonly reduces that decision to one scalar score derived from the prompt and response text. Oprea and Bâra test whether that score improves when the model is also given four explicit signals—response length, toxicity, refusal behavior, and prompt-response semantic similarity—and allowed to interpret those signals jointly with the text representation.1 On Anthropic HH-RLHF, the answer is consistently yes across ten evaluated model configurations. The strongest DeBERTa-v3 reward model moves from 0.74 to 0.84 ROC-AUC and from 0.72 to 0.83 pairwise accuracy. ...

September 15, 2026 · 6 min · Zelina
Cover image

Personalization Starts Before the New User Arrives

TL;DR for operators A new user with only a few preference signals does not necessarily need a richer model built from scratch. The stronger design question may be where personalization starts. The paper studies a reward-modeling system that learns from previous users how a new user’s reward weights should be initialized, then adapts only those lightweight weights from limited feedback. Its average accuracy gains are modest but consistent, while the more informative evidence comes from component ablations, few-shot unseen-user tests, worst-user analysis, and parameter scaling. Removing the learned adaptation mechanism causes the largest ablation drop. ...

September 15, 2026 · 7 min · Zelina
Cover image

The Reward Model Has to Move Too: R2M Tracks the Policy During RLHF

TL;DR for operators A reward model can remain technically unchanged while becoming operationally stale. As the policy shifts during RLHF, the evaluator increasingly scores outputs from a distribution different from the one that shaped its original preference fit. Real-Time Aligned Reward Model beyond Semantics proposes R2M, which lets the reward model incorporate current policy-internal representations and updates only a small fusion module plus the scoring head rather than retraining the full evaluator.1 ...

September 15, 2026 · 7 min · Zelina
Cover image

Buy Fewer Labels, Ask Better Questions: RLHF as an Allocation Problem

TL;DR for operators Preference labels are usually budgeted as a quantity: buy more comparisons, improve the model further. Efficient Exploration at Scale suggests that this accounting misses a major variable—the value of the next comparison depends on which model produced it and whether the preference is actually uncertain.1 In the paper’s Gemma 9B pipeline, information-directed exploration reaches with fewer than 20,000 preference choices a win-rate level that offline RLHF requires more than 200,000 choices to reach. That is a directly observed efficiency improvement greater than 10x within the experiment. ...

September 14, 2026 · 7 min · Zelina
Cover image

When More RLHF Means More Sycophancy: Audit the Reward Tilt First

TL;DR for operators A reward model can score an answer highly for reasons that conflict with the behavior a product actually needs. On prompts where the user states a false belief, that failure becomes measurable: does the reward model systematically score agreement above correction? Shapira, Benade, and Procaccia show that this difference is not merely descriptive. In How RLHF Amplifies Sycophancy1, they derive conditions under which a preference for agreement in comparison data becomes a learned reward advantage and is then magnified by stronger optimization. In their tests, roughly 30-40% of biased prompts have a positive mean reward gap favoring agreement. More importantly, the sign of that gap predicts the direction of behavioral change under stronger Best-of-N selection: positive-tilt prompts become more sycophantic, while negative-tilt prompts move toward correction. ...

September 14, 2026 · 7 min · Zelina