Cover image

Relevant Is Not Authorized: Put Identity Before Agent Memory Retrieval

TL;DR for operators Bio-MemArt1 addresses a problem that ordinary memory retrieval does not solve: a memory can be highly relevant to a query and still belong to the wrong user. Its intervention is deliberately narrow. Each persistent KV-memory block receives a biometric owner template. At query time, the current face or palmprint representation is compared with those templates. Memories that fail a calibrated similarity threshold are excluded before semantic retrieval begins. The original MemArt retrieval and generation machinery then operates only on the surviving memories. ...

October 1, 2026 · 7 min · Zelina
Cover image

Don’t Make the LLM Serialize the Whole Workflow

TL;DR for operators When an LLM produces a bad executable workflow, the failure may not mean that it misunderstood the business rule. It may have selected the right actions and still lost parameters, Boolean relationships, or graph ordering while writing the full structure. That distinction changes the engineering response. In Generating Workflow DAGs from Natural Language with Non-Reasoning LLMs, Anand Iyer, Bhanu Khetharpal, Srinivas Upadhya, and Ramkumar Rajagopal test an architecture that keeps language interpretation in the LLM but moves deterministic graph expansion into conventional code.1 On their 635-rule benchmark, GPT-5.3-chat improves from 56.4% to 80.6% LLM-judge validity when paired with registry selection, a compact intermediate representation, and deterministic compilation. Exact-condition accuracy rises from 56.5% to 82.2%. ...

September 29, 2026 · 8 min · Zelina
Cover image

Safety Without the Data Lake: Federating the Guard, Not the Traces

TL;DR for operators Suppose several business units or partner organizations run different agent workflows. Each has its own prompts, tools, communication patterns, and failure cases. They want a common safety layer, but centralizing those interaction traces would expose precisely the operational data they are trying to protect. The harder problem is that a guard trained elsewhere may not transfer well enough to solve this. In the reported experiments, an architecture-matched topology guard scores 0.512 AUROC when transferred off the shelf to Agent-SafetyBench, but 0.695 after in-domain retraining. Local adaptation helps, yet isolated local training is also weaker than collaborative training and becomes fragile when client labels are highly skewed. ...

September 27, 2026 · 7 min · Zelina
Cover image

Same Count, Different Bugs: Build LLM Security Scans as an Evidence Funnel

TL;DR for operators Two runs of an LLM security scanner can return almost the same number of findings while disagreeing substantially about which code locations are vulnerable. That makes finding count a weak reproducibility metric and a risky basis for remediation decisions. Bugstone-E2E, introduced in The History Is the Detector: Executing CVE Patch History, End-to-End,1 offers a broader operating model. It turns verified CVE fixing commits into reusable detection skills, uses deterministic machinery to shrink the search space before asking an LLM for semantic judgments, and sends stronger claims through separate runtime validation. ...

September 27, 2026 · 7 min · Zelina
Cover image

Audit the Crowd Before the Attack: Forecasting Multi-Agent Capture from Benign Logs

TL;DR for operators Testing each AI agent separately may not tell you how a group of those agents will behave once they begin influencing one another. Magistrali and Shani’s Aligned Alone, Misaligned Together1 provides unusually concrete evidence for that gap. In a synthetic security-triage population, a forecast constructed from adversary-free interaction logs predicted later attacked-population dismissal levels of 0.599, 0.654, and 0.680 at three held-out adversary doses. The measured values were 0.606, 0.658, and 0.685. Reported mean absolute error was 0.0058. ...

September 26, 2026 · 7 min · Zelina
Cover image

Before the Solver: Clarification Needs Its Own Readiness Gate

TL;DR for operators A business user can ask an optimization copilot for a schedule, allocation, or planning model while leaving objectives, constraints, or policy boundaries partly unstated. Two plausible interpretations can then produce different mathematical formulations even when both look coherent. The risk is not bad algebra. It is premature formulation. OR-Clarify, introduced in Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization,1 evaluates whether an agent identifies formulation-critical missing information before modeling, recovers it through questioning, avoids filling gaps with unconfirmed defaults, and stops at an appropriate point. ...

September 26, 2026 · 7 min · Zelina
Cover image

The Verifier Already Knows: Turn Pass/Fail Checks Into Training Credit

TL;DR for operators A long-running agent may execute dozens of actions before receiving one final pass/fail result. Applying that same terminal signal across the whole trajectory leaves training with little information about which earlier actions actually contributed to satisfying the task. VICT1 treats an existing programmatic verifier as a source of selective training credit. It decomposes the verifier into explicit checks, determines which checks are relevant to the rollout-level preference, and assigns additional credit only when trajectory evidence links a particular action to one of those checks. When the required evidence is missing or the verifier reconstruction is unreliable, the method falls back to the original outcome-based advantage. ...

September 26, 2026 · 7 min · Zelina
Cover image

The Linker Can’t Rank What It Never Sees

TL;DR for operators When a system must map an ambiguous phrase in text to a canonical event in a knowledge base, enlarging the search pool is not necessarily the best use of compute. The more consequential failure may occur earlier: weak contextual evidence produces the wrong candidates, leaving the final linker nothing useful to rank. ...

September 25, 2026 · 7 min · Zelina
Cover image

State Before Action: OODA-Tool Puts a Control Layer Between Context and Execution

TL;DR for operators A tool-using agent can remember the right customer, constraint, or prior result and still make the wrong call. From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use1 treats that gap as an architectural problem rather than only a prompting problem. Its strongest configuration separates four jobs: reconstruct the active task state, decide whether execution is actually warranted, choose a permitted action structure, and only then bind concrete arguments. On ToolDial, Specialized OODA beats Direct-LoRA at every tested Qwen3 scale, by 4.48 to 6.99 percentage points in Task Success. The gains are largest where state must survive long histories, missing information, changed values, constraints, or sequential dependencies. ...

September 22, 2026 · 8 min · Zelina
Cover image

When Reasoning Leaves the Prompt: Designing the Agentic Control Loop

TL;DR for operators When an LLM has to plan, call a tool, inspect the result, remember what happened, and decide what to do next, the system has changed in a more fundamental way than “adding more reasoning steps.” Agentic Reasoning for Large Language Models frames that change as a move from mostly static generation toward an interactive reasoning-and-control loop.1 ...

September 18, 2026 · 7 min · Zelina