Cover image

Storm-Chasing Agents: How EWE Turns Extreme Weather into Actionable Intelligence

Storms are easy to see after they arrive. The harder question is what actually made them happen. That distinction sounds academic until money enters the room. An insurer wants to know whether an event belongs to a changing regional risk pattern. A grid operator wants to understand whether a heatwave was driven by persistent blocking, moisture transport, or local feedback. A government agency wants a report fast enough to support preparedness, not just a polished explanation three months later. The weather event is visible. The mechanism is expensive. ...

November 28, 2025 · 14 min · Zelina
Cover image

Tile by Tile: Why LLMs Still Can't Plan Their Way Out of a 3×3 Box

A board game should not embarrass a frontier model. That is the uncomfortable charm of the 8-puzzle. It has no hidden information, no vague user intent, no messy database schema, no ambiguous policy exception, and no client saying “just make it pop.” It is a 3×3 grid with eight tiles and one blank space. Slide adjacent tiles into the blank. Reach the goal state. Done. ...

November 27, 2025 · 15 min · Zelina
Cover image

Prints Charming: How Reward Models Finally Got Serious About Long-Horizon Reasoning

Search looks simple until it becomes a workflow. A human analyst can open ten tabs, notice which source contradicts which, remember that one earlier search result changed the meaning of the question, and decide whether the next move should be another search, a calculation, or a final answer. An LLM agent can also open tabs, call tools, browse pages, run code, and produce a final answer. The difference is that the agent often does all of this with the discipline of a caffeinated intern who has been told that “more context” is the same thing as “better memory.” ...

November 25, 2025 · 13 min · Zelina
Cover image

Hierarchy, Not Hype: Why Domain Logic Beats Agent Chaos

Workflow is where agent demos go to die. A user asks for something that sounds simple: “Assess flood damage in this coastal district after the typhoon.” The agent smiles, metaphorically, and begins its little ritual. It searches, summarizes, calls a tool, thinks again, calls another tool, corrects itself, forgets one preprocessing step, invents a plausible shortcut, then produces a confident final answer that looks fine until someone who actually understands geospatial analysis asks an inconvenient question: where did the corrected satellite imagery come from? ...

November 24, 2025 · 17 min · Zelina
Cover image

Mind Over Matter: How a BDI Ontology Gives AI Agents an Actual Inner Life

Workflow agents are easy to admire until someone asks a rude but necessary question: why did the agent do that? Not “what prompt did we send?” Not “which tool did it call?” Not “can we replay the logs and hope the compliance team loses interest?” The real question is sharper: what did the agent believe, what did it want, what did it commit to doing, which plan did that commitment specify, and what evidence justified the transition from one step to the next? ...

November 24, 2025 · 18 min · Zelina
Cover image

Practice Makes Agents: How DPPO Turns Failure into Embodied Intelligence

Robots do not fail gracefully. They misread the scene, choose the wrong object, skip a physical constraint, hallucinate a plan, or produce a confident answer that would make a warehouse supervisor quietly unplug something expensive. The usual response is more data. More robot trajectories. More simulation. More web video. More carefully labelled examples. More of the industrial-scale data plumbing that makes everyone feel productive until the model still cannot decide whether a cup should be placed inside the tray or beside it. ...

November 22, 2025 · 15 min · Zelina
Cover image

Diversity Pays: Why AI Research Agents Need More Than One Good Idea

Budget has a way of making AI agents less magical. On a slide, an AI research agent looks like a neat loop: read the task, propose an idea, write code, run an experiment, improve, repeat. In production, it looks more like a slightly caffeinated junior researcher with terminal access: sometimes brilliant, sometimes stubborn, and occasionally determined to spend four hours failing at the same doomed approach because the first idea sounded respectable. ...

November 21, 2025 · 15 min · Zelina
Cover image

Game of Cones: How Physics Codes Could Fix Agent Reasoning

Controls are where agent intelligence goes to embarrass itself. Give a vision-language model a game frame, a goal, and a list of legal buttons. It may describe the scene beautifully. It may explain that the projectile is approaching, the platform is unstable, and the shiny object is probably a reward. Then it presses the wrong key, late, for the wrong duration, and walks heroically into danger. Excellent commentary. Poor organism. ...

November 21, 2025 · 16 min · Zelina
Cover image

Hex Marks the Spot: Terra Nova and the New Frontier of Agent Intelligence

A strategy game is a cruelly efficient way to embarrass an intelligent system. Not because games are magic. Not because hexagonal maps secretly contain the meaning of cognition. They do not, despite what several overexcited benchmark papers might imply after a strong coffee. Games are useful because they compress decision pressure. They make planning visible. They force trade-offs. They punish agents that confuse local competence with strategic understanding. ...

November 21, 2025 · 16 min · Zelina
Cover image

Peer Review in the Age of Agents: When Scientists Go Silicon

Reviewers are the unglamorous load-bearing wall of science. They slow things down, miss things, disagree with each other, and occasionally write comments that make authors reconsider their life choices. They are also the reason published knowledge is not just a PDF-shaped rumour. So when a conference lets AI agents act as both primary authors and reviewers, the tempting story writes itself: silicon scientists have entered the building, peer review is next, and human academics can finally retire into committee work, where they have been spiritually living for years. ...

November 21, 2025 · 16 min · Zelina