Cover image

The Verifier Already Knows: Turn Pass/Fail Checks Into Training Credit

TL;DR for operators A long-running agent may execute dozens of actions before receiving one final pass/fail result. Applying that same terminal signal across the whole trajectory leaves training with little information about which earlier actions actually contributed to satisfying the task. VICT1 treats an existing programmatic verifier as a source of selective training credit. It decomposes the verifier into explicit checks, determines which checks are relevant to the rollout-level preference, and assigns additional credit only when trajectory evidence links a particular action to one of those checks. When the required evidence is missing or the verifier reconstruction is unreliable, the method falls back to the original outcome-based advantage. ...

September 26, 2026 · 7 min · Zelina
Cover image

A Green Check Is Not a Physical Verdict: What SimVerity Changes About Agent Deployment

TL;DR for operators A simulator can clear an agent while the physical deployment simultaneously fails one operational claim and satisfies another. In the paper’s 240-trial calibration block, every source-cleared execution was a false clearance for immediate completion, while only 20 of 240 failed on reported state, 42 of 240 failed on observable physical effect, and none failed on the eventual settled postcondition. ...

September 23, 2026 · 7 min · Zelina
Cover image

The Task Is Larger Than the Prompt: What Agents Miss Before They Act

TL;DR for operators A tool-using agent can complete the action a user requested and still produce the wrong operational outcome. The missing requirement may be recoverable from current device state, a temporary preference, an accessibility setting, a privacy boundary, or the reversibility of the requested change. Implicit Intelligence – Evaluating Agents on What Users Don’t Say1 tests this problem directly. Across 205 deliberately challenging scenarios, the strongest evaluated model, GPT-5.2-pro, achieves a 48.3% Scenario Pass Rate: fewer than half of scenarios satisfy every required criterion. Its mean Normalized Scenario Score is higher, at 72.7%, showing that agents often complete substantial parts of the task while still missing at least one consequential requirement. ...

September 8, 2026 · 8 min · Zelina
Cover image

Reasoning Under a Running Clock: Why Agent Rankings Reverse in Real Time

TL;DR for operators A better plan can become a worse agent when the environment keeps moving while the system thinks. STAR makes that reversal unusually clear: Kimi-K2-Thinking leads the unlimited-deliberation evaluation with a rating of 1206.1 and a 1.00 win rate, then falls to 842.6 and a 0.210 win rate in real-time play, where GLM-4.6 leads at 1180.8.1 ...

September 7, 2026 · 7 min · Zelina
Cover image

Success Is Not the System: Rethinking How AI Agents Should Be Evaluated

TL;DR for operators An enterprise agent can finish a workflow and still be a poor production system. It may require repeated retries, call the wrong tool before recovering, exceed an acceptable cost envelope, fail under small environmental changes, or depend on a human to stop a consequential action. Bin Xu’s survey, AI Agent Systems: Architectures, Applications, and Evaluation, treats those behaviors as part of the system being evaluated, not as incidental implementation details.1 Its central abstraction places the model inside an execution loop with memory, tools, verifiers, and an environment. Section 6 then evaluates the resulting system across multiple dimensions rather than collapsing performance into task success. ...

September 7, 2026 · 5 min · Zelina
Cover image

Progress Is Not Completion: What CAP Reveals About Browser-Agent Readiness

TL;DR for operators A browser automation can search several sites, manipulate interfaces, gather useful information, and return a polished response while still missing one requirement that makes the workflow unusable. It may apply the wrong filter, misread a value in a chart, or fail to notice that a panel is collapsed. For production decisions, visible progress is not the same thing as reliable completion. ...

August 30, 2026 · 8 min · Zelina
Cover image

Same Agent, Different Audience: When Social Pressure Changes the Recommendation

TL;DR for operators A company may deploy an AI adviser whose recommendation is visible to a sponsor, manager, funding partner, or future evaluator. Even when the task, model, assigned role, and public interaction history remain matched, changing who can see the answer—and what that audience may control—can substantially change the recommendation. The study compares two responses generated by the same agent at the same point in the interaction: one visible to the consequential audience and one framed as confidential. It compares changes in decisions, reasoning, and consistency across the two channels rather than treating either response as the agent’s true belief. For the targeted agent, decision divergence increased from 2.8% at baseline to 39.9% under relationships that made alignment socially advantageous, while the untargeted control agent remained comparatively stable. ...

July 30, 2026 · 7 min · Zelina
Cover image

The Skill Library That Could Read but Couldn’t Run

TL;DR for operators A system can discover agent “skills” that look coherent to humans and still fail to make an agent more capable. That is the useful result of Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining.1 The authors build a pipeline that cuts GUI interaction histories into segments, clusters those segments into candidate routines, generates explicit skill specifications, and trains a Qwen3-8B policy with Group Relative Policy Optimization, or GRPO, to compose the resulting skills. ...

July 15, 2026 · 21 min · Zelina
Cover image

Skill Issue, Literally: Repairing Agent Instructions Without an Answer Key

TL;DR for operators Runbooks decay. APIs shift, data schemas mutate, file paths move, and the “expert procedure” that worked last quarter starts quietly steering an agent into a wall. The paper behind this article, SkillAudit: Ground-Truth-Free Skill Evolution via Paired Trajectory Auditing, asks a useful operational question: can an agent skill be improved when nobody has provided hidden tests, reference answers, task rewards, or expert labels?1 ...

July 3, 2026 · 21 min · Zelina
Cover image

Clawing Back the Benchmark: When AI Agents Start Testing Themselves

Tickets. That is where the future of AI agents becomes less theatrical and more irritatingly real. Not in a glossy demo where an agent books a holiday after three polite prompts, but in a helpdesk queue where it must read a ticket, check a knowledge base, update a CRM record, avoid leaking private data, recover from a failed API call, and still produce something a human manager can audit later. ...

April 23, 2026 · 17 min · Zelina