Cover image

Buy Fewer Labels, Ask Better Questions: RLHF as an Allocation Problem

TL;DR for operators Preference labels are usually budgeted as a quantity: buy more comparisons, improve the model further. Efficient Exploration at Scale suggests that this accounting misses a major variable—the value of the next comparison depends on which model produced it and whether the preference is actually uncertain.1 In the paper’s Gemma 9B pipeline, information-directed exploration reaches with fewer than 20,000 preference choices a win-rate level that offline RLHF requires more than 200,000 choices to reach. That is a directly observed efficiency improvement greater than 10x within the experiment. ...

September 14, 2026 · 7 min · Zelina
Cover image

Preference Signals, Not Preference Theater

Preference Signals, Not Preference Theater Businesses are currently learning an expensive lesson: user behavior is not the same thing as user preference. A person clicks because the button was large. A driver brakes because the situation was unclear. A customer accepts a chatbot answer because the refund is small and arguing is tedious. A manager approves a workflow because the dashboard made the alternative invisible. The log file looks objective. It is also quietly contaminated by habit, uncertainty, exploration, friction, fatigue, and the occasional human desire to end the meeting before lunch. ...

June 3, 2026 · 15 min · Zelina
Cover image

Mind the Reward Gap: Why Business AI Needs More Than Pretty Answers

Opening — Why this matters now Business AI has entered its awkward teenage years. The first phase was easy to admire: models could draft, summarize, classify, recommend, and explain. Then companies started asking the rude adult questions: Can we trust the answer? Did it make the right trade-off? Can it improve from outcomes? What happens when the reward signal is wrong? ...

May 2, 2026 · 17 min · Zelina