Cover image

Alignment Is a Coverage Problem Before It Is a Loss-Function Problem

TL;DR for operators An alignment team with a fixed preference dataset faces a deceptively simple decision: train offline with a direct method, or spend more compute to keep generating and evaluating new responses during training. The cheaper route is not always the safer one. The survey by Tarun Raheja and Nilay Pochhi1 highlights a theoretical coverage result under which offline contrastive preference learning needs stronger coverage of possible responses than online reinforcement learning. If useful responses lie outside the regions represented in the fixed dataset, an offline learner has no direct learning signal there. Online methods can generate new data and therefore operate under a weaker, partial-coverage requirement. ...

September 15, 2026 · 8 min · Zelina