Backtrack to Breakthrough: Why Great AI Agents Revisit
GSM-Agent shows that agent quality depends less on raw search time than on whether the agent knows when to return to promising evidence with a better query.
GSM-Agent shows that agent quality depends less on raw search time than on whether the agent knows when to return to promising evidence with a better query.
UltraHorizon shows why long-horizon AI agents fail less like weak chatbots and more like badly managed investigation teams.
EELMA shows how agent optionality can become useful telemetry—but only when we distinguish real control from noisy behavioural diversity.
A mechanism-first reading of why reinforcement learning helps LLM planning through exploration, why policy-gradient can collapse into brittle one-path behaviour, and why Q-learning only helps when rewards expose the structure of the task.
Shachi shows why serious LLM-based agent simulation needs modular cognitive architecture, not just better personas and larger crowds.
How SaMuLe turns failed agent traces into a reusable diagnostic layer—and what that means for enterprise automation.
A practical reading of CORE, a path-based evaluation framework showing why tool-using AI agents must be judged by the sequence of actions they take, not only the state they leave behind.
A mechanism-first look at why reasoning traces can make AI agents both harder to fool and better at fooling each other.
Recon-Act shows why browser agents may improve less by clicking harder and more by converting repeated failures into governed, reusable tools.
A practical reading of what task-free LLM agents do when nobody gives them a job: they build, self-test, or disappear into recursive self-description.