Cover image

Think Again, but Make It Count: Train Reflection Before You Spend More Tokens on It

TL;DR for operators A reasoning model can spend extra tokens reconsidering its answer and still fail to repair the mistake. Worse, unnecessary reconsideration can disturb reasoning that was already correct. The paper studied here treats that as a training problem rather than an invitation to add another inference-time review loop. It first filters a model’s own critiques using verifiable ground truth, then trains the model on the surviving examples. A second reinforcement-learning stage rewards both final-answer correctness and the quality of the reflective step itself. ...

September 16, 2026 · 7 min · Zelina
Cover image

Failures, Taxonomized: How Multi‑Level Reflection Turns Agents Into Self‑Learners

Failure is usually treated as waste. The demo breaks, the agent apologises, someone adds a prompt patch, and everyone pretends the next retry will be more mature. Very enterprise. Very ceremonial. The SaMuLe paper makes a more useful claim: failed agent runs are not just embarrassing logs. They are the curriculum.1 More precisely, they are raw material for a structured reflection pipeline that turns messy trajectories into error taxonomies, cross-task lessons, and finally a small retrospective model trained to diagnose future failures. ...

October 2, 2025 · 14 min · Zelina