Personalization Starts Before the New User Arrives
TL;DR for operators A new user with only a few preference signals does not necessarily need a richer model built from scratch. The stronger design question may be where personalization starts. The paper studies a reward-modeling system that learns from previous users how a new user’s reward weights should be initialized, then adapts only those lightweight weights from limited feedback. Its average accuracy gains are modest but consistent, while the more informative evidence comes from component ablations, few-shot unseen-user tests, worst-user analysis, and parameter scaling. Removing the learned adaptation mechanism causes the largest ablation drop. ...