Cover image

Think Again, but Make It Count: Train Reflection Before You Spend More Tokens on It

TL;DR for operators A reasoning model can spend extra tokens reconsidering its answer and still fail to repair the mistake. Worse, unnecessary reconsideration can disturb reasoning that was already correct. The paper studied here treats that as a training problem rather than an invitation to add another inference-time review loop. It first filters a model’s own critiques using verifiable ground truth, then trains the model on the surviving examples. A second reinforcement-learning stage rewards both final-answer correctness and the quality of the reflective step itself. ...

September 16, 2026 · 7 min · Zelina
Cover image

More Critics, Less Gain: Self-Questioning Has a Stability Limit

TL;DR for operators When a model can check its own reasoning, more self-checks are not automatically better. On GSM8K, Llama-3.2-1B rises from a 33.14% chain-of-thought baseline to 35.28% with one alternative critique and 35.84% with two, but falls back to 33.43% with three. The broader analysis links higher disagreement among these self-generated alternatives to greater reward variance and less stable policy updates. ...

September 15, 2026 · 6 min · Zelina
Cover image

Talking to Yourself, but Make It Useful: Intrinsic Self‑Critique in LLM Planning

“Please double-check your work” is one of the least expensive quality-control systems ever invented. It is also one of the least dependable. A person who overlooked a constraint the first time may overlook it again. A language model is no different, except that it can produce a longer and more persuasive explanation of why the overlooked constraint was never important. ...

January 3, 2026 · 17 min · Zelina
Cover image

Stepwise Think-Critique: Teaching LLMs to Doubt Themselves (Productively)

The useful part of doubt is timing Doubt is not useful after the invoice is paid, the client report is sent, or the model has already produced a confident wrong answer with twelve decorative paragraphs of reasoning. At that point, “let us verify” becomes less like quality control and more like archaeology. ...

December 18, 2025 · 16 min · Zelina