Cover image

A Green Check Is Not a Physical Verdict: What SimVerity Changes About Agent Deployment

TL;DR for operators A simulator can clear an agent while the physical deployment simultaneously fails one operational claim and satisfies another. In the paper’s 240-trial calibration block, every source-cleared execution was a false clearance for immediate completion, while only 20 of 240 failed on reported state, 42 of 240 failed on observable physical effect, and none failed on the eventual settled postcondition. ...

September 23, 2026 · 7 min · Zelina
Cover image

State Before Action: OODA-Tool Puts a Control Layer Between Context and Execution

TL;DR for operators A tool-using agent can remember the right customer, constraint, or prior result and still make the wrong call. From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use1 treats that gap as an architectural problem rather than only a prompting problem. Its strongest configuration separates four jobs: reconstruct the active task state, decide whether execution is actually warranted, choose a permitted action structure, and only then bind concrete arguments. On ToolDial, Specialized OODA beats Direct-LoRA at every tested Qwen3 scale, by 4.48 to 6.99 percentage points in Task Success. The gains are largest where state must survive long histories, missing information, changed values, constraints, or sequential dependencies. ...

September 22, 2026 · 8 min · Zelina
Cover image

More Compute, Different Jobs: Choosing Inference-Time Reliability Controls

TL;DR for operators When one model answer can trigger a consequential downstream step, spending more inference compute before trusting that answer can help—but the form of that spending matters. In the reported experiment, generating several reasoning attempts and aggregating the answer that repeatedly emerged increased verified acceptance from 56.2% to 64.9%. Asking the same model to critique and revise itself produced a smaller increase, from 47.2% to 50.6%. Adding a second model did not improve the reported acceptance rate: 47.4% of outputs survived cross-model verification versus a 48.7% single-model baseline. ...

September 17, 2026 · 7 min · Zelina
Cover image

One Step Is Not a Workflow: Where LLM Rule Following Starts to Break

TL;DR for operators A model that is highly reliable at applying one explicit rule transition is not necessarily reliable at executing an entire procedure built from those transitions. In Reasoning Capabilities of Large Language Models. Lessons Learned from General Game Playing1, the strongest evaluated model, Gemini 2.5 Pro, achieves 95.6% exact success on one-step next-state generation. At five dependent state transitions, exact success falls to 73.4%. When the model must also choose actions during those five steps, it falls again to 65.3%. ...

September 17, 2026 · 7 min · Zelina
Cover image

The Fourth Hop Changes the Risk Profile: Measuring Reliability in Multi-Step LLM Workflows

TL;DR for operators When one model output becomes input to the next stage, a final accuracy score tells you too little about where reliability is being lost. A workflow may fail because a required fact was never available, because a later composition step is intrinsically harder, or because an earlier mistake was allowed to propagate. Those failure modes call for different controls. ...

September 17, 2026 · 8 min · Zelina
Cover image

An 8/10 Is Not a Probability: Validating LLM Confidence Before It Controls Workflow

TL;DR for operators A model saying “8/10 confident” does not mean its underlying uncertainty is approximately 20%. Across the evaluated settings, the average instance-level correlation between reported confidence and logits-based confidence is only 0.135. The more useful operating rule is narrower. First test whether reported scores vary enough to distinguish cases. Then measure whether those scores rank examples meaningfully on held-out data. Separately test whether their numerical scale agrees with the comparison signal and whether they are calibrated against correctness. Do not substitute one test for another. ...

September 10, 2026 · 7 min · Zelina
Cover image

Are You Sure? Reliability Starts After the First Answer

TL;DR for operators A model can answer correctly, be challenged by the user, and then talk itself into being wrong. Deployment gates that look only at first-turn accuracy or reported confidence can miss that failure. Saadat and Nemzer’s Certainty Robustness Benchmark1 tests 200 LiveBench math and reasoning questions with independent follow-ups: “Are you sure?”, “You are wrong!”, and a request for 1–100 confidence. GPT-5.2 and Claude Sonnet 4.5 began at almost the same accuracy, yet each collapsed under a different form of pushback. ...

September 10, 2026 · 5 min · Zelina
Cover image

What to Commit First: IGFD Turns Token Order Into a Reliability Lever

TL;DR for operators When a model has several unresolved positions, committing the token it predicts most confidently is not necessarily the best use of that commitment. A predictable punctuation mark may add little information, while a semantic token can make several nearby predictions easier. Information-Guided Frontier Decoding (IGFD) changes that choice. Fang et al.1 rank candidate commitments using the token’s own confidence, uncertainty in neighboring unresolved positions, a penalty for structural tokens, and a locality constraint on where commitments can occur. ...

September 9, 2026 · 8 min · Zelina
Cover image

The Task Is Larger Than the Prompt: What Agents Miss Before They Act

TL;DR for operators A tool-using agent can complete the action a user requested and still produce the wrong operational outcome. The missing requirement may be recoverable from current device state, a temporary preference, an accessibility setting, a privacy boundary, or the reversibility of the requested change. Implicit Intelligence – Evaluating Agents on What Users Don’t Say1 tests this problem directly. Across 205 deliberately challenging scenarios, the strongest evaluated model, GPT-5.2-pro, achieves a 48.3% Scenario Pass Rate: fewer than half of scenarios satisfy every required criterion. Its mean Normalized Scenario Score is higher, at 72.7%, showing that agents often complete substantial parts of the task while still missing at least one consequential requirement. ...

September 8, 2026 · 8 min · Zelina
Cover image

When Hallucination Is More Than a Wrong Fact: Measuring Reliability Through the User

TL;DR for operators A model change can improve an automated hallucination benchmark while leaving users dissatisfied for a different reason: sources are hard to verify, reasoning appears unsupported, false claims are stated with confidence, or corrections are ignored. The System Hallucination Scale (SHS) gives teams a structured way to measure those experiences across five dimensions rather than reducing reliability to a binary factual-error judgment.1 ...

September 8, 2026 · 7 min · Zelina