Cover image

Relevant Is Not Authorized: Put Identity Before Agent Memory Retrieval

TL;DR for operators Bio-MemArt1 addresses a problem that ordinary memory retrieval does not solve: a memory can be highly relevant to a query and still belong to the wrong user. Its intervention is deliberately narrow. Each persistent KV-memory block receives a biometric owner template. At query time, the current face or palmprint representation is compared with those templates. Memories that fail a calibrated similarity threshold are excluded before semantic retrieval begins. The original MemArt retrieval and generation machinery then operates only on the surviving memories. ...

October 1, 2026 · 7 min · Zelina
Cover image

More Memory, Worse Decisions: Why Agent Recall Needs Routing

TL;DR for operators An agent has a recurring allocation problem: how much of its past should it bring back before answering a question or choosing its next action? More history increases the chance that useful evidence is available, but it also increases the amount of material competing with the information that matters now. ...

August 27, 2026 · 8 min · Zelina
Cover image

After the Bad Memory: Repairing the Decisions It Already Touched

TL;DR for operators When an agent discovers that a stored customer preference, prior observation, or workflow fact was wrong, deleting that record may be too late. The faulty information may already have shaped a plan, triggered a tool call, entered the final answer, or created new persistent memories. Yu et al. propose a repair mechanism that follows those dependencies rather than resetting everything.1 On their 150-case controlled benchmark, it recovered 85.3% of cases, compared with 77.3% for LLM-judge repair, while reducing the replay ratio from 21.7% to 12.3% and average LLM calls from 9.80 to 5.70. ...

August 21, 2026 · 7 min · Zelina
Cover image

The Memory Score Changed Before the Memory Did

TL;DR for operators A team should be able to replace one memory component, compare two vendors, or combine user history with multimodal records without rebuilding the entire agent stack. But a benchmark score does not measure the memory algorithm alone. It also reflects when memory is formed and retrieved, how updates accumulate, and whether connected components exchange the fields each one expects. ...

August 3, 2026 · 9 min · Zelina
Cover image

Agree Once, Remember Later: The Commit Boundary in Personal Agents

TL;DR for operators A personal assistant can hear a confident user claim, store it as a preference or rule, and rely on it during a later task after the original conversation is gone. The safety problem is therefore not only the agreeable reply. It is the write that lets the claim survive. ...

July 24, 2026 · 8 min · Zelina
Cover image

Memory Has to Earn Its Keep

TL;DR for operators Memory is not valuable because an agent writes something down. That is called logging. Sometimes it is called “reflection,” if the logging has better branding. The paper Enhancing Software Engineering Through Closed-Loop Memory Optimization introduces MemOp, a framework for software-engineering agents that defines memory utility by downstream impact: a memory is useful only if it improves the agent’s later performance on software tasks.1 The important move is not the existence of Memory.md, nor the idea that past trajectories can be summarized. The important move is the loop: generate memory from an agent trajectory, validate whether that memory improves task performance, reject harmful or redundant memories, and train a memory model using the resulting accepted and rejected examples. ...

June 27, 2026 · 17 min · Zelina
Cover image

Memory Lane Meets Mainframe: Why Coding Agents Need Better Memories, Not Bigger Egos

Memory is a familiar word. That is exactly why it can mislead us. When people hear that coding agents need “memory,” the first image is often a giant scrapbook: past prompts, previous patches, command logs, successful code snippets, failed attempts, and whatever else the agent has dragged behind it like a very confident intern with a messy backpack. More memory sounds safer. More traces sound more useful. More remembered work sounds like less repeated work. ...

April 16, 2026 · 17 min · Zelina
Cover image

Thinking Fast, Remembering Slow: Why SWE-AGILE Fixes the Memory Crisis of AI Agents

Memory sounds like a storage problem. Give the agent a longer context window, let it keep the full conversation, and the work should become easier. This is the kind of solution that looks obvious until it meets a real software repository, a failing test suite, a long terminal log, and a model that now has to find one important clue buried somewhere in the middle of its own autobiography. ...

April 14, 2026 · 18 min · Zelina
Cover image

Memory, Rewritten: Why ByteRover Kills the Pipeline (and Maybe Saves Agents)

The agent did not forget. The system outsourced remembering. Memory sounds like a solved engineering problem until an agent has to use it for work. A customer-support agent remembers the refund policy but not why an exception was approved. A research agent retrieves the right document but loses the reasoning trail that connected three earlier notes. A workflow agent crashes halfway through a task, comes back online, and must reconstruct its own state from search results like a detective investigating a crime it personally committed. ...

April 5, 2026 · 18 min · Zelina
Cover image

Autonomous Memory: When AI Starts Debugging Itself

Memory sounds glamorous until someone has to maintain it. In a demo, memory is easy. The agent remembers your name, recalls your last project, and maybe retrieves that one document you uploaded three sessions ago. Very charming. Very investor-deck friendly. Then the system goes into production. The memory store grows. Similar events blur together. Image captions lose details. Timestamps drift. Retrieval starts pulling almost-right context. The model becomes confidently nostalgic about things that did not happen. ...

April 2, 2026 · 21 min · Zelina