Buy Fewer Labels, Ask Better Questions: RLHF as an Allocation Problem
TL;DR for operators Preference labels are usually budgeted as a quantity: buy more comparisons, improve the model further. Efficient Exploration at Scale suggests that this accounting misses a major variable—the value of the next comparison depends on which model produced it and whether the preference is actually uncertain.1 In the paper’s Gemma 9B pipeline, information-directed exploration reaches with fewer than 20,000 preference choices a win-rate level that offline RLHF requires more than 200,000 choices to reach. That is a directly observed efficiency improvement greater than 10x within the experiment. ...