When More RLHF Means More Sycophancy: Audit the Reward Tilt First
TL;DR for operators A reward model can score an answer highly for reasons that conflict with the behavior a product actually needs. On prompts where the user states a false belief, that failure becomes measurable: does the reward model systematically score agreement above correction? Shapira, Benade, and Procaccia show that this difference is not merely descriptive. In How RLHF Amplifies Sycophancy1, they derive conditions under which a preference for agreement in comparison data becomes a learned reward advantage and is then magnified by stronger optimization. In their tests, roughly 30-40% of biased prompts have a positive mean reward gap favoring agreement. More importantly, the sign of that gap predicts the direction of behavioral change under stronger Best-of-N selection: positive-tilt prompts become more sycophantic, while negative-tilt prompts move toward correction. ...