Cover image

From Chains to Trees: Why LLM Agents Need Structural Memory

A mechanism-first reading of T-STAR, showing why multi-turn LLM agents learn better when failed and successful rollouts are compared as shared trees rather than isolated chains.

April 9, 2026 · 18 min · Zelina
Cover image

The Map Is Not the Territory—But Your LLM Thinks It Is

EVGeoQA shows why tool-using LLM agents still struggle with real-world spatial planning: they can reason locally, but often fail to explore enough.

April 9, 2026 · 16 min · Zelina
Cover image

The Memory Isn’t the Point — It’s the Feeling: Why AI Needs Affective Memory, Not Just Recall

A-MBER shows why long-term AI assistants need selective, structured affective memory—not just larger context windows—to understand what users feel now.

April 9, 2026 · 17 min · Zelina
Cover image

The Minimal LLM Thesis: When Agents Think for Themselves

A decomposition study shows why agent performance may come from measurable harness structure before it comes from larger or more frequent LLM calls.

April 9, 2026 · 14 min · Zelina
Cover image

Unsolvable by Design: Turning AI Plans Into Security Guarantees

A mechanism-first reading of planning task shielding: how AI planning can be used to make dangerous states unreachable, where the guarantee holds, and where the computation breaks.

April 9, 2026 · 16 min · Zelina
Cover image

When Feelings Negotiate: Why Emotion Might Be the Missing Layer in AI Agents

A mechanism-first reading of EmoMAS and what strategic emotional orchestration means for business-facing AI agents.

April 9, 2026 · 18 min · Zelina
Cover image

Benchmarking the Benchmarks: Why ACE-Bench Might Be the Missing Layer in Agent Evaluation

A mechanism-first reading of AgentCE-Bench, showing why controllable agent evaluation may be more useful than another realism-heavy leaderboard.

April 8, 2026 · 14 min · Zelina
Cover image

Blinded by Design: When AI Stops Thinking and Starts Remembering

A practical reading of epistemic blinding: an inference-time audit protocol for separating LLM reasoning from memorized entity priors in business-critical ranking workflows.

April 8, 2026 · 19 min · Zelina
Cover image

Claw-Eval — When Agents Game the System, the System Needs Claws

Claw-Eval shows why serious AI-agent evaluation must audit behavior, stress-test recovery, and separate lucky success from deployable reliability.

April 8, 2026 · 16 min · Zelina
Cover image

From Spreadsheets to Swarms: How Agentic AI Rewrites the Retail Supply Chain

A mechanism-first reading of Flowr, an agentic AI framework that turns supermarket replenishment from manual coordination into supervised workflow automation.

April 8, 2026 · 18 min · Zelina