Cover image

When the Reward Is Right but the Incentive Is Wrong

TL;DR for operators A preference score can rank responses correctly and still be the wrong signal to feed directly into an alignment system. Wang et al. show why: when deployment deliberately keeps the aligned model close to its base behavior, the base model’s own response probabilities continue to influence what gets generated. Their Stackelberg Reward Shaping framework therefore changes the reward landscape rather than simply increasing reward strength.1 ...

September 14, 2026 · 8 min · Zelina