Cover image

When the Reward Is Right but the Incentive Is Wrong

TL;DR for operators A preference score can rank responses correctly and still be the wrong signal to feed directly into an alignment system. Wang et al. show why: when deployment deliberately keeps the aligned model close to its base behavior, the base model’s own response probabilities continue to influence what gets generated. Their Stackelberg Reward Shaping framework therefore changes the reward landscape rather than simply increasing reward strength.1 ...

September 14, 2026 · 8 min · Zelina
Cover image

Shape the Signal, Keep the Objective: MeRLa’s Bet on Reusable RLHF Rewards

TL;DR for operators RLHF teams usually have two obvious levers: improve the reward model or improve the policy optimizer. Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback introduces a third.1 MeRLa learns an additional task-aware reward signal across auxiliary tasks, freezes it, and adds it to the existing reward during subsequent policy optimization. ...

August 23, 2026 · 7 min · Zelina