Cover image

Reasoning Under a Running Clock: Why Agent Rankings Reverse in Real Time

TL;DR for operators A better plan can become a worse agent when the environment keeps moving while the system thinks. STAR makes that reversal unusually clear: Kimi-K2-Thinking leads the unlimited-deliberation evaluation with a rating of 1206.1 and a 1.00 win rate, then falls to 842.6 and a 0.210 win rate in real-time play, where GLM-4.6 leads at 1180.8.1 ...

September 7, 2026 · 7 min · Zelina
Cover image

When Model Output Can Change State: An Architecture Guide to Agent Reliability

TL;DR for operators A tool-using agent does more than generate an answer. It observes part of a workflow, carries information forward, decides what to do, changes external state, and then reacts to the result. A wrong answer in a chatbot may remain text; a wrong action in an agent can alter a file, submit a transaction, call the wrong service, or create a bad state that later decisions treat as valid. ...

September 6, 2026 · 7 min · Zelina
Cover image

Train the Agent Before the Sandbox Exists

TL;DR for operators A team can have stable API contracts before it has an executable sandbox, populated user state, and reliable integration infrastructure. The usual assumption is that serious multi-step agent training must wait, because later API responses need to reflect what earlier calls changed. ESAT challenges that sequencing. Lee et al. generate stateful training trajectories from API specifications without the target environment, then validate and filter the interactions before training.1 On AppWorld, training with ESAT-S52 trajectories generated from 52 synthetic applications independent of AppWorld improves Task Goal Completion by 8.4 to 47.0 percentage points across evaluated models and usually outperforms training on trajectories collected inside the real AppWorld environment. ...

September 3, 2026 · 7 min · Zelina
Cover image

More Memory, Worse Decisions: Why Agent Recall Needs Routing

TL;DR for operators An agent has a recurring allocation problem: how much of its past should it bring back before answering a question or choosing its next action? More history increases the chance that useful evidence is available, but it also increases the amount of material competing with the information that matters now. ...

August 27, 2026 · 8 min · Zelina
Cover image

Decision Rights, Not More Layers: What an Auditable Fraud Pipeline Actually Earns

TL;DR for operators A fraud classifier has already scored a transaction. The next operational choice is whether extra context—relationship patterns, anomaly signals, explanations, or an LLM investigator—should merely inform the case or be allowed to change the decision. In Rahil Sharma’s evaluation, the answer is component-specific.1 The bounded LLM investigator was correct on 39 of 60 deliberately balanced difficult cases, versus 43 of 60 for simply applying a 0.5 threshold to the classifier: 65.0% versus 71.7%. The agent changed eight classifier decisions. Two changes fixed mistakes; six replaced correct decisions with incorrect ones. ...

August 24, 2026 · 7 min · Zelina
Cover image

Higher Pass Rate, More Broken Tasks: The Regression Tax in Agent Skill Libraries

TL;DR for operators Adding reusable instructions to an agent creates a release-management problem that average accuracy does not fully expose. A library can solve tasks the baseline missed while simultaneously breaking tasks the baseline handled correctly. Darshan Tank and Baran Nama measure that trade-off directly in The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents.1 Across 18 skill-library conditions, they observe 553 baseline-fail-to-skill-pass transitions but also 324 baseline-pass-to-skill-fail transitions. The new failures offset 59% of the gross gains. ...

August 19, 2026 · 7 min · Zelina
Cover image

The Trace Has the Answer, Not the Alternative: Agentic-DPO for Offline Agent Training

TL;DR for operators Historical agent traces usually tell you what a successful operator or agent did. They do not tell you which plausible alternative the current model is most likely to choose incorrectly. That missing contrast is the problem Agentic-DPO targets.1 Instead of sending the student through full environment rollouts, Agentic-DPO pauses at states already present in expert trajectories, samples several one-step actions from the current student, and contrasts the expert action with a plausible different action the student actually favors. On StableToolBench with Qwen3.5-2B, plain SFT reaches 57.1% canonical accuracy while Agentic-DPO reaches 90.9%. ...

August 19, 2026 · 8 min · Zelina
Cover image

The Catalog Grew. The Agent Needed a Call Stack.

TL;DR for operators As an agent’s tool catalog grows, it must solve two linked problems: choosing the right capability without carrying every tool schema into each decision, and remembering where to return after several nested actions. The paper’s hierarchy addresses both by showing the model only the options relevant to its current branch and storing nested workflow state explicitly. ...

August 3, 2026 · 9 min · Zelina
Cover image

Same Agent, Different Audience: When Social Pressure Changes the Recommendation

TL;DR for operators A company may deploy an AI adviser whose recommendation is visible to a sponsor, manager, funding partner, or future evaluator. Even when the task, model, assigned role, and public interaction history remain matched, changing who can see the answer—and what that audience may control—can substantially change the recommendation. The study compares two responses generated by the same agent at the same point in the interaction: one visible to the consequential audience and one framed as confidential. It compares changes in decisions, reasoning, and consistency across the two channels rather than treating either response as the agent’s true belief. For the targeted agent, decision divergence increased from 2.8% at baseline to 39.9% under relationships that made alignment socially advantageous, while the untargeted control agent remained comparatively stable. ...

July 30, 2026 · 7 min · Zelina
Cover image

Refusal Is Not a Result: Vera Tests What Agents Actually Changed

TL;DR for operators A production agent can refuse a dangerous request after its tools have already changed a repository, sent a message, or altered an account. That is why the final response alone cannot establish whether the system behaved safely: stated refusal, attempted action, and persistent environmental change may point to different conclusions. ...

July 29, 2026 · 8 min · Zelina