When More Problems Stop Helping: RL Data Scaling Becomes an Allocation Problem
TL;DR for operators A post-training team with a fixed compute budget has several ways to spend it: add more verified problems, generate variants of existing problems, increase difficulty, or expose the model to structurally different tasks. The usual accounting metric—number of training problems—does not tell the team which choice produces the most useful learning experience. ...