Reward the Recheck: Reflection as a Control Surface for Reasoning Post-Training
A mathematical-reasoning study shows why reward design and training-stage compatibility can matter more than simply adding another post-training step.
A mathematical-reasoning study shows why reward design and training-stage compatibility can matter more than simply adding another post-training step.
PRoSFI shows how machine-checkable intermediate reasoning can raise measured soundness without requiring a language model to generate full formal proofs.
A mathematical-reasoning study suggests that supervising the decision that determines a solution path can outperform training on the entire reasoning trace while using far fewer tokens.
A reasoning-centered framework for deciding when an AI workflow needs tools, persistent adaptation, multi-agent coordination, or post-training rather than more prompting.
Multilingual reasoning systems need trace-quality signals that track useful reasoning, not merely resemblance to English.
A controlled comparison shows that sampling, cross-model checking, and self-critique spend inference compute on different reliability problems—and should not be treated as interchangeable.
A Theory of Mind study shows why inference-time reasoning should be routed by task complexity and answer format rather than treated as a default quality upgrade.
A formal game benchmark shows why strong one-step rule following does not automatically justify autonomous multistep execution.
ROI-Reasoning shows that under a shared inference budget, deciding where not to spend reasoning effort can matter as much as increasing reasoning depth.
Omanic shows why multi-step LLM reliability requires separating missing knowledge, later-step composition difficulty, and propagated upstream errors.