Think Again, but Make It Count: Train Reflection Before You Spend More Tokens on It
TL;DR for operators A reasoning model can spend extra tokens reconsidering its answer and still fail to repair the mistake. Worse, unnecessary reconsideration can disturb reasoning that was already correct. The paper studied here treats that as a training problem rather than an invitation to add another inference-time review loop. It first filters a model’s own critiques using verifiable ground truth, then trains the model on the surviving examples. A second reinforcement-learning stage rewards both final-answer correctness and the quality of the reflective step itself. ...