Cover image

The File System Strikes Back: Why AI Agents Still Can’t Understand Your Life

Files are where AI agent demos go to become adults. In a product video, the agent opens a few clean documents, remembers your preferences, drafts an answer, books the meeting, and looks quietly inevitable. In an actual computer, the same agent faces a folder called final_final_v3, a receipt saved as an image, a calendar invite with the wrong title, a video that contains the decisive evidence at second 8, and three people who all appear in the same user’s digital life. Suddenly the assistant that “knows you” looks less like a colleague and more like an intern who has discovered search for the first time. ...

April 2, 2026 · 17 min · Zelina
Cover image

Friction Over Fiction: Why AI Agents Need to Feel Resistance

Tools are not free. That sentence sounds too obvious to deserve an article, which is usually a warning that the industry has built several architectures pretending it is false. A tool-using AI agent can call a search API, query a database, inspect a document, ask another model, trigger a diagnostic pipeline, or run a workflow step. In a clean demo, each call feels like another harmless unit of intelligence. The agent thinks, acts, observes, thinks again, and the audience applauds because the trace looks busy. Busy is often mistaken for capable. Enterprise software has enjoyed this little confusion for decades. ...

April 1, 2026 · 17 min · Zelina
Cover image

Blueprints for Thinking: Why CAD Needs Agents, Not Prompts

A bracket looks simple until someone has to manufacture it. On a screen, a generated part can look almost right: the flange appears round, the bolt holes seem evenly spaced, and the central bore is visible enough to satisfy a casual glance. Then a machinist opens the file, measures it, and discovers the inconvenient details: the wall thickness is wrong, a boolean cut failed, two solids merely touch instead of joining, or the bounding box is off by a few millimeters. ...

March 30, 2026 · 17 min · Zelina
Cover image

From Blueprints to Prompts: Automating Building–Grid Intelligence with LLM Agents

Building simulation is not glamorous work. It is a room full of configuration files, simulator interfaces, reward functions, time-series outputs, and small mistakes that quietly invalidate a week of analysis. The industry likes to talk about intelligent buildings. The less marketable truth is that before a building can be intelligent, someone has to wire the experiment together correctly. ...

March 30, 2026 · 16 min · Zelina
Cover image

The Parallel Mind: How AIRA2 Turns AI Research from Guesswork into Scalable Discovery

Research has a waiting-room problem. A human team proposes an experiment, waits for the training run, checks the metric, argues about whether the result is real, then decides what to try next. The cycle is familiar, expensive, and mildly theatrical. AI research agents promise to compress that loop. Give the agent a benchmark, a compute budget, and a tool environment; let it search; harvest better models at the end. Convenient. Also, if done naively, a beautiful machine for producing confident nonsense at GPU speed. ...

March 30, 2026 · 18 min · Zelina
Cover image

ARC-AGI-3 — When AI Stops Guessing and Starts Thinking

Demo days are generous. A sales engineer opens a prepared workflow, the agent clicks through a familiar sequence, the dashboard turns green, and everyone politely pretends not to notice how much of the intelligence was smuggled into the setup. ARC-AGI-3 is less polite. The paper introduces an interactive benchmark for agentic intelligence: not a static puzzle, not a multiple-choice exam, and not a coding task with a unit test waiting like a benevolent parent. An agent enters a novel, abstract, turn-based environment. It receives no explicit objective. It must explore, infer the rules, identify what counts as success, build a working model of the environment, and execute a plan efficiently.1 ...

March 28, 2026 · 16 min · Zelina
Cover image

Driving by Words: When LLMs Take the Wheel (Literally)

Taxi. That is the easiest way to understand the paper. Not because Vega is a robotaxi system. It is not. But because a taxi ride exposes the missing layer in many autonomous-driving discussions: the passenger does not merely want the car to obey traffic rules. The passenger wants the car to behave under intent. ...

March 28, 2026 · 14 min · Zelina
Cover image

Harnessing the Harness: When AI Stops Being a Model Problem

Glue is not glamorous. In most AI product discussions, the model gets the spotlight. The harness—the scripts, prompts, validators, retry rules, state files, tool adapters, and stopping criteria around the model—gets treated as plumbing. Necessary, slightly annoying, and best ignored until it leaks. That habit is becoming expensive. The paper Natural-Language Agent Harnesses argues that the surrounding execution system is no longer a secondary implementation detail. It is often the actual unit of agent performance, reliability, and portability.1 The paper’s useful claim is not that “natural language replaces code.” That would be a lovely fantasy for people who have not debugged parsers, sandboxes, or file permissions lately. The sharper claim is that part of the harness can become an editable natural-language policy object, while exact execution remains in code. ...

March 28, 2026 · 16 min · Zelina
Cover image

Agent Factories: When More AI Means Better Hardware

Button. That was the promise of High-Level Synthesis: write a high-level program, push it through the toolchain, and receive efficient hardware without spending the afternoon whispering to pragmas like a medieval engineer negotiating with silicon spirits. The button never quite arrived. HLS did raise the abstraction level from RTL to C/C++. But performance still depends on expert choices: where to pipeline, where not to pipeline, which arrays to partition, which loops to unroll, which memory access pattern is quietly sabotaging the whole design. The code looks like software; the reasoning remains hardware. ...

March 27, 2026 · 14 min · Zelina
Cover image

EcoThink: When AI Learns to Think Less (and Achieve More)

A chatbot does not need a philosophy seminar to answer “Who directed Oppenheimer?” That sentence sounds obvious. Yet a large part of today’s AI infrastructure behaves as if every user query deserves a carefully staged internal drama: retrieve facts, reason through them, verify the logic, produce a chain of intermediate steps, and finally deliver the answer the system could have produced with a simple lookup. It is impressive in the same way using a crane to move a coffee cup is impressive. Technically capable. Operationally absurd. ...

March 27, 2026 · 14 min · Zelina