From Chains to Trees: Why LLM Agents Need Structural Memory
A mechanism-first reading of T-STAR, showing why multi-turn LLM agents learn better when failed and successful rollouts are compared as shared trees rather than isolated chains.
A mechanism-first reading of T-STAR, showing why multi-turn LLM agents learn better when failed and successful rollouts are compared as shared trees rather than isolated chains.
EVGeoQA shows why tool-using LLM agents still struggle with real-world spatial planning: they can reason locally, but often fail to explore enough.
A-MBER shows why long-term AI assistants need selective, structured affective memory—not just larger context windows—to understand what users feel now.
A decomposition study shows why agent performance may come from measurable harness structure before it comes from larger or more frequent LLM calls.
A mechanism-first reading of planning task shielding: how AI planning can be used to make dangerous states unreachable, where the guarantee holds, and where the computation breaks.
A mechanism-first reading of EmoMAS and what strategic emotional orchestration means for business-facing AI agents.
A mechanism-first reading of AgentCE-Bench, showing why controllable agent evaluation may be more useful than another realism-heavy leaderboard.
A practical reading of epistemic blinding: an inference-time audit protocol for separating LLM reasoning from memorized entity priors in business-critical ranking workflows.
Claw-Eval shows why serious AI-agent evaluation must audit behavior, stress-test recovery, and separate lucky success from deployable reliability.
A mechanism-first reading of Flowr, an agentic AI framework that turns supermarket replenishment from manual coordination into supervised workflow automation.