Cover image

The Planner Trusted the Wrong State: A New Security Boundary for Embodied Agents

TL;DR for operators An embodied agent can receive the correct user instruction and still plan toward the wrong objective if the internal description of its environment has been manipulated. Liu et al. test this failure mode by altering planner-visible state semantics rather than changing the instruction, model, planner, executor, or environment itself.1 ...

September 6, 2026 · 7 min · Zelina
Cover image

When Model Output Can Change State: An Architecture Guide to Agent Reliability

TL;DR for operators A tool-using agent does more than generate an answer. It observes part of a workflow, carries information forward, decides what to do, changes external state, and then reacts to the result. A wrong answer in a chatbot may remain text; a wrong action in an agent can alter a file, submit a transaction, call the wrong service, or create a bad state that later decisions treat as valid. ...

September 6, 2026 · 7 min · Zelina
Cover image

Audit the Step, Not the Aftershock

TL;DR for operators When an autonomous analysis agent makes one questionable operation, every later operation may inherit the damage. An audit that merely asks which steps look unusual can therefore produce a long queue of downstream symptoms rather than isolate where new violations occurred. Ahmed Hassoon and Mark Dredze formalize a different target: score each operation according to whether it behaved as expected given the state it actually received.1 Under their assumptions, a correctly executed downstream step remains statistically null even when its input was already corrupted. This makes one-step scoring useful for narrowing a review queue. ...

August 28, 2026 · 8 min · Zelina
Cover image

After the Bad Memory: Repairing the Decisions It Already Touched

TL;DR for operators When an agent discovers that a stored customer preference, prior observation, or workflow fact was wrong, deleting that record may be too late. The faulty information may already have shaped a plan, triggered a tool call, entered the final answer, or created new persistent memories. Yu et al. propose a repair mechanism that follows those dependencies rather than resetting everything.1 On their 150-case controlled benchmark, it recovered 85.3% of cases, compared with 77.3% for LLM-judge repair, while reducing the replay ratio from 21.7% to 12.3% and average LLM calls from 9.80 to 5.70. ...

August 21, 2026 · 7 min · Zelina
Cover image

Wait, Let Me Check: Why Long-CoT AI Can Still Verify the Wrong Thing

Checking is supposed to calm people down. In business, a second review makes a financial model feel safer. A compliance checklist makes a release feel governed. A senior analyst saying “let me double-check that” gives the room a small dopamine hit of procedural seriousness. Long Chain-of-Thought models have learned the same theatre. They pause. They reconsider. They say “wait.” They verify arithmetic. They sometimes generate reasoning traces so long that one begins to feel the model must be thinking deeply, if only because wasting that many tokens while being shallow seems rude. ...

June 9, 2026 · 19 min · Zelina
Cover image

Protocol Over Hype: Why AI Drug Discovery Agents Need Memory, Not Just Models

Drug discovery is a wonderful place for AI demos. The model proposes a molecule, the molecule looks plausible, a docking score improves, and the slide deck starts to glow with that familiar color: almost-commercial blue. Then the evaluation protocol arrives and ruins the party. The problem is simple, and therefore easy to underestimate. A drug discovery agent is rarely asked to return one impressive molecule. It is asked to return a set of molecules that jointly satisfies several requirements: enough candidates, enough diversity, acceptable binding proxies, drug-likeness, synthetic accessibility, novelty, and other threshold-style constraints. One molecule can look good. A few molecules can look good. The final returned pool can still fail. ...

April 13, 2026 · 15 min · Zelina
Cover image

Verify Before You Automate: Why AI Agents Need an Internal Audit Function

A number is a small thing. One integer in one answer. A seating capacity, a contract limit, a delivery quantity, a tax threshold, a credit exposure. Nothing dramatic. Certainly not the sort of thing that should become an architecture problem. Then an AI agent guesses it, sounds confident, stores the guess, and uses it again later. ...

April 10, 2026 · 18 min · Zelina
Cover image

Middleware Matters: Why Your AI Agent Needs a Lifecycle (Not Just a Brain)

Agent demos are easy to like because nothing important is attached to them. A demo agent can call the wrong tool, misread a JSON response, or politely announce that an API failure is actually a useful answer. Everyone smiles, someone says “interesting,” and the team adds another item to the backlog. Very innovative. Very safe. Very far from production. ...

March 17, 2026 · 19 min · Zelina
Cover image

Mirror, Mirror on the Agent: Teaching LLMs to Judge Their Own Actions

The agent did exactly what it was taught. That was the problem. A familiar business agent failure does not look dramatic. It looks boring. The agent searches the database, clicks the wrong record, receives an error, retries the same action, receives the same error, retries again, and then politely informs the user that it has encountered “temporary difficulty.” Very professional. Completely useless. ...

March 12, 2026 · 16 min · Zelina
Cover image

Consistency Is Not a Coincidence: When LLM Agents Disagree With Themselves

A support ticket arrives. The agent reads the same customer history, sees the same policy document, and has access to the same tools. On Monday, it searches for the refund rule, retrieves the correct clause, and gives a clean answer. On Tuesday, with the same input, it searches for a different phrase, retrieves a less relevant document, wanders through two extra steps, and ends with a confident answer that is only approximately useful. ...

February 14, 2026 · 16 min · Zelina