When Reasoning Pays (and When It Cheats): Fixing RL Signals in LLM Training
A mechanism-first reading of PAPO, showing why separating correctness rewards from process rubrics can keep reasoning-model RL useful without paying models to perform for the judge.