Yesterday’s Gold, Today’s Bias: Reusing Human Feedback After a Model Upgrade
TL;DR for operators An archive of expensive human corrections does not automatically remain a valid training target after the production model improves. In the paper introducing StalePO,1 legacy post-edits are often closer to the older NMT system than to the upgraded model, yet they can still contain corrections worth recovering. StalePO handles that mismatch by learning the relative signal in the old preference pair while explicitly protecting the new model’s existing response and constraining changes at token level. ...