Cover image

Reward the Recheck: Reflection as a Control Surface for Reasoning Post-Training

A mathematical-reasoning study shows why reward design and training-stage compatibility can matter more than simply adding another post-training step.

September 18, 2026 · 7 min · Zelina
Cover image

The Right Answer Is Not a Proof: Put Verification Inside the Reasoning Loop

PRoSFI shows how machine-checkable intermediate reasoning can raise measured soundness without requiring a language model to generate full formal proofs.

September 18, 2026 · 7 min · Zelina
Cover image

Train the Decision, Not the Transcript: FSLR Targets the First Reasoning Choice

A mathematical-reasoning study suggests that supervising the decision that determines a solution path can outperform training on the entire reasoning trace while using far fewer tokens.

September 18, 2026 · 6 min · Zelina
Cover image

When Reasoning Leaves the Prompt: Designing the Agentic Control Loop

A reasoning-centered framework for deciding when an AI workflow needs tools, persistent adaptation, multi-agent coordination, or post-training rather than more prompting.

September 18, 2026 · 7 min · Zelina
Cover image

English-Like Is Not the Same as Correct: Measuring Multilingual Reasoning Quality

Multilingual reasoning systems need trace-quality signals that track useful reasoning, not merely resemblance to English.

September 17, 2026 · 7 min · Zelina
Cover image

More Compute, Different Jobs: Choosing Inference-Time Reliability Controls

A controlled comparison shows that sampling, cross-model checking, and self-critique spend inference compute on different reliability problems—and should not be treated as interchangeable.

September 17, 2026 · 7 min · Zelina
Cover image

More Thought Is Not Always More Reliable: Routing Reasoning for Social Judgment

A Theory of Mind study shows why inference-time reasoning should be routed by task complexity and answer format rather than treated as a default quality upgrade.

September 17, 2026 · 7 min · Zelina
Cover image

One Step Is Not a Workflow: Where LLM Rule Following Starts to Break

A formal game benchmark shows why strong one-step rule following does not automatically justify autonomous multistep execution.

September 17, 2026 · 7 min · Zelina
Cover image

Reasoning Has an Opportunity Cost: Allocate Compute Before You Spend It

ROI-Reasoning shows that under a shared inference budget, deciding where not to spend reasoning effort can matter as much as increasing reasoning depth.

September 17, 2026 · 6 min · Zelina
Cover image

The Fourth Hop Changes the Risk Profile: Measuring Reliability in Multi-Step LLM Workflows

Omanic shows why multi-step LLM reliability requires separating missing knowledge, later-step composition difficulty, and propagated upstream errors.

September 17, 2026 · 8 min · Zelina