Your Reward Model Is Also a Voting Rule: What Social Choice Changes About RLHF
TL;DR for operators When thousands of annotators disagree about two acceptable model responses, the pipeline still has to decide how that disagreement becomes model behavior. That decision is often hidden inside familiar technical choices: how comparisons are sampled, whether votes are collapsed into majority labels, which reward-model class is fitted, and how aggressively a policy is optimized against the resulting score. ...