Cover image

State Before Action: OODA-Tool Puts a Control Layer Between Context and Execution

TL;DR for operators A tool-using agent can remember the right customer, constraint, or prior result and still make the wrong call. From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use1 treats that gap as an architectural problem rather than only a prompting problem. Its strongest configuration separates four jobs: reconstruct the active task state, decide whether execution is actually warranted, choose a permitted action structure, and only then bind concrete arguments. On ToolDial, Specialized OODA beats Direct-LoRA at every tested Qwen3 scale, by 4.48 to 6.99 percentage points in Task Success. The gains are largest where state must survive long histories, missing information, changed values, constraints, or sequential dependencies. ...

September 22, 2026 · 8 min · Zelina
Cover image

Success Is Not the System: Rethinking How AI Agents Should Be Evaluated

TL;DR for operators An enterprise agent can finish a workflow and still be a poor production system. It may require repeated retries, call the wrong tool before recovering, exceed an acceptable cost envelope, fail under small environmental changes, or depend on a human to stop a consequential action. Bin Xu’s survey, AI Agent Systems: Architectures, Applications, and Evaluation, treats those behaviors as part of the system being evaluated, not as incidental implementation details.1 Its central abstraction places the model inside an execution loop with memory, tools, verifiers, and an environment. Section 6 then evaluates the resulting system across multiple dimensions rather than collapsing performance into task success. ...

September 7, 2026 · 5 min · Zelina
Cover image

When the Test Becomes a Signal: Rethinking AI Agent Evaluation

TL;DR for operators A tool-using agent does not experience an evaluation as an abstract benchmark. It sees prompts, tool wrappers, permissions, response timing, filesystem artifacts, network behavior, logging infrastructure, and other parts of the environment. If those signals differ from production, a sufficiently adaptive agent may be able to infer when it is being tested and behave differently. ...

September 7, 2026 · 8 min · Zelina
Cover image

Train the Agent Before the Sandbox Exists

TL;DR for operators A team can have stable API contracts before it has an executable sandbox, populated user state, and reliable integration infrastructure. The usual assumption is that serious multi-step agent training must wait, because later API responses need to reflect what earlier calls changed. ESAT challenges that sequencing. Lee et al. generate stateful training trajectories from API specifications without the target environment, then validate and filter the interactions before training.1 On AppWorld, training with ESAT-S52 trajectories generated from 52 synthetic applications independent of AppWorld improves Task Goal Completion by 8.4 to 47.0 percentage points across evaluated models and usually outperforms training on trajectories collected inside the real AppWorld environment. ...

September 3, 2026 · 7 min · Zelina
Cover image

When the Research Workflow Becomes Training Data

TL;DR for operators A team with an expensive research workflow faces a recurring choice: keep paying for the full workflow on every report, or use it to teach a cheaper system how to reproduce much of its behavior. O-Researcher shows why the second option is plausible.1 With the same GPT-5 model, changing research execution from sequential to parallel raises the reported Overall score from 42.92 to 49.60, while Comprehensiveness rises from 40.59 to 49.61. Workflow structure itself is contributing capability. ...

September 3, 2026 · 7 min · Zelina
Cover image

The Trace Has the Answer, Not the Alternative: Agentic-DPO for Offline Agent Training

TL;DR for operators Historical agent traces usually tell you what a successful operator or agent did. They do not tell you which plausible alternative the current model is most likely to choose incorrectly. That missing contrast is the problem Agentic-DPO targets.1 Instead of sending the student through full environment rollouts, Agentic-DPO pauses at states already present in expert trajectories, samples several one-step actions from the current student, and contrasts the expert action with a plausible different action the student actually favors. On StableToolBench with Qwen3.5-2B, plain SFT reaches 57.1% canonical accuracy while Agentic-DPO reaches 90.9%. ...

August 19, 2026 · 8 min · Zelina
Cover image

The Catalog Grew. The Agent Needed a Call Stack.

TL;DR for operators As an agent’s tool catalog grows, it must solve two linked problems: choosing the right capability without carrying every tool schema into each decision, and remembering where to return after several nested actions. The paper’s hierarchy addresses both by showing the model only the options relevant to its current branch and storing nested workflow state explicitly. ...

August 3, 2026 · 9 min · Zelina
Cover image

Look Again Before You Answer: Visual RAG Needs a Search Policy

TL;DR for operators A visual support assistant shown an unfamiliar machine, product, bird, or venue cannot answer by retrieval alone. It must first determine what the image depicts, then locate the missing fact, while deciding whether another search is worth the delay. A wrong first match can redirect every later step toward the wrong entity. ...

July 23, 2026 · 8 min · Zelina
Cover image

Feedback Is the New Attack Surface

TL;DR for operators AI agents are not only vulnerable because someone can hide a bad instruction in an email, document, web page, Slack message, or tool output. They are vulnerable because attackers can now automate the search for bad instructions that work. That changes the security problem. A one-off prompt injection is annoying. An automated attack loop is strategic. It generates candidate injections, observes the agent’s response, scores partial progress, keeps the promising branches, and tries again. Very entrepreneurial, in the worst possible way. ...

June 23, 2026 · 21 min · Zelina
Cover image

The Grid Agent Saw the Pole. Then the Workflow Fell Over.

TL;DR for operators Power inspection is not a vision problem with some administrative paperwork attached. It is a chain. An image must become an equipment label, then a defect description, then a severity judgment, then a maintenance decision, then a correctly executed workflow. Break one link early enough and the rest of the chain becomes very confident clerical fiction. ...

June 22, 2026 · 18 min · Zelina