Cover image

DeltaEvolve: When Evolution Learns Its Own Momentum

Memory is usually where agentic systems go to become expensive. That is not the glamorous failure mode. It is not the cinematic robot rebellion, nor the slightly more realistic spreadsheet full of hallucinated invoices. It is quieter: an LLM agent keeps improving a program, stores previous attempts, retrieves a few “good” ones, and then spends half its context window rereading code scaffolding that no longer explains anything useful. ...

February 5, 2026 · 16 min · Zelina
Cover image

Perspective Without Rewards: When AI Develops a Point of View

AI agents do not need feelings to become difficult to read. That is already enough trouble. A long-running agent can enter a workflow, absorb context, make decisions, and gradually behave as though the situation has a particular “shape.” The system may not merely react to the latest input. It may carry forward a learned orientation: this client is risky, this process is stable, this market regime is noisy, this user wants speed more than precision. In ordinary product language, we call that “context.” In engineering dashboards, we often reduce it to memory, state, embeddings, or hidden activations. In philosophical language, one might be tempted to call it a perspective. ...

February 5, 2026 · 14 min · Zelina
Cover image

Conducting the Agents: Why AORCHESTRA Treats Sub-Agents as Recipes, Not Roles

Agent teams are easy to draw and hard to run. On a slide, the architecture looks comforting: a planner, a researcher, a coder, a reviewer, perhaps a compliance agent standing in the corner with a clipboard. Everyone has a role. Everyone collaborates. The diagram is tidy, which is usually the first warning sign. ...

February 4, 2026 · 14 min · Zelina
Cover image

Search-R2: When Retrieval Learns to Admit It Was Wrong

Search is supposed to make language models safer. The model does not know something, so it searches. It finds evidence, reasons over that evidence, and gives a better answer. Very civilized. Very responsible. Then the first search query goes slightly wrong. The model retrieves a relevant-looking but misleading paragraph. It builds the next reasoning step around the wrong entity. The next query becomes narrower, but in the wrong direction. The final answer may still sound fluent, because fluency is the one department where language models rarely file sick leave. The actual reasoning chain, however, has already drifted. ...

February 4, 2026 · 16 min · Zelina
Cover image

When Your Agent Starts Copying Itself: Breaking Conversational Inertia

A support agent keeps asking the same diagnostic question after the customer has already answered it. A research agent revisits the same failed source path with slightly different wording. A workflow agent tries the same invalid action again because, apparently, the best evidence for what to do next is what it just did badly. ...

February 4, 2026 · 17 min · Zelina
Cover image

DRIFT-BENCH: When Agents Stop Asking and Start Breaking

A user says, “Update the record with a sensible value.” That sentence is small. The damage may not be. For a normal chatbot, the worst outcome might be a vague answer wearing a confident expression. Annoying, yes, but usually recoverable. For an agent connected to a database, file system, workflow platform, or API service, the same ambiguity becomes operational. The model may update the wrong row, call the wrong endpoint, overwrite a file, or politely explain its mistake after making it. Charming, in the same way a self-driving forklift is charming. ...

February 3, 2026 · 17 min · Zelina
Cover image

Seeing Is Not Reasoning: Why Mental Imagery Still Breaks Multimodal AI

A model can generate a pretty sequence of images. Good. So can a slide deck. The harder question is whether those images actually help it think. That is the uncomfortable point behind MentisOculi: Revealing the Limits of Reasoning with Mental Imagery, a new benchmark paper that tests whether frontier multimodal models can do something closer to human mental imagery: form a visual state, keep it stable, transform it step by step, and use the transformed state to decide what to do next.1 Not merely “look at an image and answer a question.” Not “draw a plausible intermediate picture.” Actual visual reasoning, with consequences. ...

February 3, 2026 · 18 min · Zelina
Cover image

When LLMs Meet Time: Why Time-Series Reasoning Is Still Hard

Dashboard numbers are seductive because they look obedient. Revenue goes up, traffic dips, latency spikes, inventory turns over, temperature drifts, volatility clusters. Put the sequence into a chart and the pattern seems almost polite. Then someone asks an LLM what happened. The model answers fluently. It may even sound like an analyst who has seen too many quarterly review decks and has developed a protective layer of confidence. But fluency is not temporal understanding. A model can describe a curve, name a trend, and still fail to understand which segment comes next, whether a transformation is correct, or whether a discontinuity is an error or a legitimate feature of the process. ...

February 3, 2026 · 16 min · Zelina
Cover image

FadeMem: When AI Learns to Forget on Purpose

Memory is easy to sell. Give an AI agent a bigger context window. Add a vector database. Store every user preference, meeting note, support ticket, and half-correct instruction that ever passed through the system. Then call it “persistent memory,” because apparently a drawer full of old receipts is now intelligence. The problem is that agents do not fail only because they forget. They also fail because they remember too much, too flatly, and too obediently. Old facts compete with new ones. Repeated but trivial details crowd out rare but important constraints. Retrieval brings back something semantically similar but temporally wrong. The agent sounds confident because the database found something. Very helpful. Very dangerous. ...

February 1, 2026 · 13 min · Zelina
Cover image

When Empathy Needs a Map: Benchmarking Tool‑Augmented Emotional Support

Empathy is easy to fake for one sentence. A chatbot can say “that sounds exhausting” without knowing anything about you, your situation, your city, your time zone, or whether the advice it is about to give is physically possible. That is the awkward part of emotional support AI: the tone can be soft while the facts are made of air. A very caring assistant can still recommend a midnight walk at 3 p.m., suggest a closed café, or confidently invent local details because it wants to be helpful. The kindness is real enough in style. The grounding is not. ...

February 1, 2026 · 16 min · Zelina