When the Wrong Label Shouts Loudest: Correcting Preference Noise in RLHF and DPO
TL;DR for operators Preference pipelines have an awkward failure mode: when an annotator accidentally chooses the worse of two responses, a standard preference loss can push hardest on exactly the pair the model ranks most strongly against the recorded label. A bad label can therefore receive unusually strong corrective force instead of being naturally ignored. ...