Cover image

Agents, Automata, and the Memory of Thought

A booking agent is not dangerous because it can “reason.” It is dangerous because it can remember the wrong thing, forget the right thing, loop politely forever, or book the flight before the human has actually confirmed. The philosophy department may enjoy debating whether this counts as intention. The operations team has a simpler question: can we know, before deployment, what behaviours this system can produce? ...

November 1, 2025 · 15 min · Zelina
Cover image

The Benchmark Awakens: AstaBench and the New Standard for Agentic Science

Procurement meetings have a habit of turning AI agents into theatre. A vendor shows a polished research assistant. It finds papers, writes a summary, cites sources, maybe generates a small experiment plan. Everyone nods. Someone says “agentic workflow.” Someone else says “autonomous discovery.” A budget appears. The machine is declared practically scientific, which is convenient, because the machine itself has not yet been asked to survive the boring parts of science: retrieval under controlled conditions, code execution, data analysis, experimental reproduction, hypothesis testing, and the small matter of completing all required steps without wandering into the digital bushes. ...

October 31, 2025 · 13 min · Zelina

From Field Notes to Farm Operating Intelligence

A high-value commercial farm redesigned daily crop, irrigation, pest, harvest, labor, and buyer-delivery coordination around a reviewed AI operations brief instead of fragmented messages and manager memory.

October 30, 2025 · 8 min · Vox
Cover image

Beyond Utility: When LLM Agents Start Dreaming Their Own Tasks

A task list is usually where enterprise automation becomes reassuringly boring. Someone defines the work. The system executes it. A dashboard turns green, or, in more honest organisations, amber with an explanation. The point is not mystery. The point is control. The paper behind this article, LLM Agents Beyond Utility: An Open-Ended Perspective, asks what happens when that tidy arrangement is disturbed: what if the agent does not merely complete tasks, but proposes them? What if it can remember what it has done, inspect its environment, write notes to itself, and continue across runs?1 ...

October 23, 2025 · 15 min · Zelina
Cover image

Blueprints of Agency: Compositional Machines and the New Architecture of Intelligence

A prototype begins innocently enough: a product team wants a small machine, a vehicle, a tool, a fixture, perhaps a mechanism that throws something across a room because medieval engineering apparently never left the group chat. The modern AI pitch says the agent can design it. Give it parts, constraints, and a goal; let it reason; let it test; let it improve. ...

October 23, 2025 · 14 min · Zelina
Cover image

When Lateral Beats Linear: How LToT Rethinks the Tree of Thought

Budget is easy to approve when the system still fails anyway. That is the awkward little problem sitting underneath many agentic AI roadmaps. A product team adds more inference tokens, more retries, more tool calls, more reflective loops, and more polite internal monologue. The demo becomes slower, the invoice becomes more interesting, and the model still sometimes walks straight past the right answer because it pruned the wrong branch three steps ago. Progress, apparently. ...

October 21, 2025 · 13 min · Zelina

From Claim Chaos to Review-Ready Case Files

A small insurance broker redesigned a fragmented claims-preparation workflow into a human-reviewed agentic process that turns scattered documents into completeness-checked, risk-screened, underwriter-ready files.

October 15, 2025 · 9 min · Vox
Cover image

The Mr. Magoo Problem: When AI Agents 'Just Do It'

Office automation has a simple seduction: give the agent a task, let it click through the mess, and reclaim the human hours previously sacrificed to forms, folders, email threads, and software that looks as if it was last loved in 2009. That is the promise. The problem is that some agents take the phrase “complete the task” a little too personally. ...

October 9, 2025 · 17 min · Zelina
Cover image

Backtrack to Breakthrough: Why Great AI Agents Revisit

Search is easy. Knowing when to go back is harder. That is the useful irritation inside GSM-Agent, a new benchmark for studying agentic reasoning under controlled conditions.1 The paper takes grade-school maths problems from GSM8K, removes the premises from the prompt, hides those premises in a searchable document database, and asks an LLM agent to recover the facts before solving the problem. The arithmetic is not supposed to be impressive. That is the point. If a model fails here, we cannot calmly blame differential geometry, PhD-level law, or some mysteriously adversarial enterprise workflow. The agent simply did not find and use the facts. ...

October 3, 2025 · 15 min · Zelina
Cover image

Options = Power: Turning Empowerment into a KPI for AI Agents

Login. That is where many agent evaluations become strangely unserious. A benchmark asks whether the agent completed a task. A dashboard records whether the browser session ended successfully. A monitoring system checks whether the tool call returned an error. Then the agent enters valid credentials and suddenly gains access to a much larger part of the environment. ...

October 3, 2025 · 16 min · Zelina