Cover image

More Compute, Different Jobs: Choosing Inference-Time Reliability Controls

TL;DR for operators When one model answer can trigger a consequential downstream step, spending more inference compute before trusting that answer can help—but the form of that spending matters. In the reported experiment, generating several reasoning attempts and aggregating the answer that repeatedly emerged increased verified acceptance from 56.2% to 64.9%. Asking the same model to critique and revise itself produced a smaller increase, from 47.2% to 50.6%. Adding a second model did not improve the reported acceptance rate: 47.4% of outputs survived cross-model verification versus a 48.7% single-model baseline. ...

September 17, 2026 · 7 min · Zelina
Cover image

One Step Is Not a Workflow: Where LLM Rule Following Starts to Break

TL;DR for operators A model that is highly reliable at applying one explicit rule transition is not necessarily reliable at executing an entire procedure built from those transitions. In Reasoning Capabilities of Large Language Models. Lessons Learned from General Game Playing1, the strongest evaluated model, Gemini 2.5 Pro, achieves 95.6% exact success on one-step next-state generation. At five dependent state transitions, exact success falls to 73.4%. When the model must also choose actions during those five steps, it falls again to 65.3%. ...

September 17, 2026 · 7 min · Zelina
Cover image

The Wrong Answer May Start Before Reasoning

TL;DR for operators A model reads a chart, diagram, or photographed math problem and produces the wrong answer. Treating that event as a generic “reasoning failure” can send engineering effort to the wrong component. The model may have misread a number, attached a label to the wrong object, confused a scale or unit, or reasoned incorrectly after extracting the right facts. ...

September 13, 2026 · 8 min · Zelina
Cover image

Spend Verification Where Risk Is Highest

TL;DR for operators A long generated report can contain low-risk statements alongside claims that deserve much stronger evidence checking. Applying the strictest verification rule everywhere spends verifier capacity without distinguishing where factual failure is most likely. FACTOR turns that problem into a routing decision. The framework estimates uncertainty for individual claims, applies progressively stricter evidence requirements as estimated risk rises, and then selects among multiple generated candidates. In the reported benchmark, FACTOR reached a FActScore of 42.3 versus 36.8 for static verification while reducing average verification calls from 194.0 to 41.9.1 ...

September 11, 2026 · 7 min · Zelina
Cover image

Higher Pass Rate, More Broken Tasks: The Regression Tax in Agent Skill Libraries

TL;DR for operators Adding reusable instructions to an agent creates a release-management problem that average accuracy does not fully expose. A library can solve tasks the baseline missed while simultaneously breaking tasks the baseline handled correctly. Darshan Tank and Baran Nama measure that trade-off directly in The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents.1 Across 18 skill-library conditions, they observe 553 baseline-fail-to-skill-pass transitions but also 324 baseline-pass-to-skill-fail transitions. The new failures offset 59% of the gross gains. ...

August 19, 2026 · 7 min · Zelina
Cover image

The Proof Is in the Process

TL;DR for operators MaxProof is not primarily a story about a model suddenly becoming brilliant at mathematics. It is a story about wrapping an imperfect model in a disciplined production process. MiniMax trains M3 to perform three distinct jobs: write proofs, identify concrete errors in proofs, and repair proofs using those critiques. At inference time, MaxProof generates a population of candidate solutions, evaluates them conservatively, preserves competing approaches, applies both targeted patches and broader rewrites, and finally chooses one answer through pairwise comparison.1 ...

July 10, 2026 · 19 min · Zelina
Cover image

Do the Math, Not the Mime: Why LLM Reasoning Needs a Verification Pipeline

A spreadsheet error rarely announces itself with dramatic music. It usually arrives politely. A pricing model gives a clean answer. A compliance calculator writes a confident explanation. A financial assistant produces a neat derivation with enough intermediate steps to look reassuring. The result is formatted, fluent, and possibly wrong. That is the uncomfortable business lesson behind Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges, a 2026 survey of roughly 120 studies on LLM mathematical reasoning.1 The paper is not introducing one new benchmark, one heroic model, or one more leaderboard trophy to place on the already overcrowded mantelpiece. Its useful contribution is more structural: it connects datasets, representations, training methods, tool use, verifiers, and evaluation metrics into one reasoning pipeline. ...

May 31, 2026 · 14 min · Zelina
Cover image

Do the Math, Not the Mime: Why LLM Reasoning Needs a Verification Pipeline

Spreadsheet errors have a special talent: they look boring until they become expensive. That is the business version of the LLM math problem. A model can produce a calm, step-by-step explanation, put a confident number at the bottom, and still be wrong in the only place that matters. Worse, the reasoning may look plausible enough that a manager, analyst, tutor, or compliance reviewer nods and moves on. The answer has the rhythm of thinking. It has the costume of calculation. It may even have a chain-of-thought trace. Very civilized. Still not proof. ...

May 30, 2026 · 19 min · Zelina
Cover image

The Proof Is in the Instance: Why AI Safety Can’t Be Fully Verified

The verifier that cannot know everything Verification sounds like the sensible adult in the AI safety room. The model may hallucinate, the benchmark may flatter, the demo may sparkle under conference lighting, but the verifier is supposed to be the hard stop: a formal mechanism that checks whether an AI system’s behavior satisfies a specified policy. ...

April 7, 2026 · 17 min · Zelina
Cover image

FAME or Fortune? How Formal Explanations Finally Scale to Real Neural Networks

Audit is a boring word until the model says something expensive. A credit model rejects an applicant. A visual inspection model flags a component. A traffic-sign classifier keeps its prediction under small pixel changes. The business question is not merely, “What did the model look at?” That is the demo-room version. The operational question is harder: which input features must remain fixed so that the model’s decision is guaranteed not to change under allowed perturbations? ...

March 13, 2026 · 16 min · Zelina