Cover image

Double Helix, Double Checks: Why Agentic AI Needs Governance Before It Writes Your Code

Code is where AI confidence goes to become expensive. A chatbot can produce a plausible function in ten seconds. An agent can now plan a refactor, split files, update interfaces, generate documentation, and politely leave behind a system that fails because one event payload forgot a required field. Very efficient. Very modern. Very annoying. ...

March 5, 2026 · 16 min · Zelina
Cover image

From Prompt Chains to Algebra: Why Agentics 2.0 Treats AI Workflows Like Math

Workflow diagrams lie. They make AI systems look orderly: one box extracts information, another box reasons, a third box writes a conclusion, and a final box sends the result somewhere official-looking. In production, of course, the boxes often exchange blobs of fragile text, half-structured JSON, hidden assumptions, and one optimistic prompt that begins with “You are an expert…” ...

March 5, 2026 · 15 min · Zelina
Cover image

Agents in the Lab: When Bayesian Adversaries Keep AI Scientists Honest

Lab work has an old rule: never trust the first beautiful result. It may be correct. It may also be a measurement artifact wearing a lab coat. That rule becomes more important when the “research assistant” is an LLM that can write code, invent tests, explain errors, and occasionally hallucinate with the confidence of a junior consultant who has just discovered PowerPoint. The paper “AI-for-Science Low-code Platform with Bayesian Adversarial Multi-Agent Framework” takes this problem seriously.1 Its central claim is not that scientific automation needs a larger model, a longer prompt, or another cheerful agent named “Planner.” The claim is sharper: in AI-assisted scientific coding, both the generated code and the generated tests are uncertain. If the validator is also an LLM, then the system has not solved hallucination. It has merely hired hallucination as compliance staff. ...

March 4, 2026 · 15 min · Zelina
Cover image

Mind the Gap: Why Agency Isn’t Intelligence (Yet)

A trading bot keeps executing while the market regime changes. A warehouse robot keeps optimizing its route while a sensor slowly drifts. A customer-service agent keeps sounding fluent while the conversation loses coherence one turn at a time. From the outside, the system still looks agentic. It acts. It responds. It may even keep producing acceptable short-term outcomes. The dashboard, naturally, waits until the mess is obvious. Dashboards are polite like that. ...

February 28, 2026 · 16 min · Zelina
Cover image

Mirror, Mirror on the LLM: Teaching Models to Think About Their Thinking

Evidence is not the same as judgment. Anyone who has watched an AI assistant work through a multi-document question has seen the strange version of this failure. The model finds the relevant fact. It even says something that looks like the right answer. Then, a few paragraphs later, it invents an extra condition, follows that condition with great confidence, and lands somewhere else. ...

February 28, 2026 · 15 min · Zelina
Cover image

Attention with Doubt: Teaching Transformers When *Not* to Trust Themselves

Confidence is cheap. A classifier can always give you a probability. The awkward question is whether that probability deserves to be believed. This is not a philosophical problem when the model is recommending a movie. It becomes expensive when the model is screening documents, triaging support tickets, flagging fraud, routing legal clauses, or deciding whether a case should be escalated to a human. In those settings, “92% confident” is not decoration. It is an operating instruction. ...

February 5, 2026 · 16 min · Zelina
Cover image

When LLMs Lose the Plot: Diagnosing Reasoning Instability at Inference Time

Mistakes are easy to audit after the fact. That is why most AI evaluation still behaves like a mildly disappointed teacher: wait for the final answer, mark it right or wrong, and pretend the interesting part happened at the end. But in real LLM workflows, the damage often starts earlier. A model begins with a plausible line of reasoning, then drifts. It changes route without noticing. It over-explains a wrong intermediate step. It doubles back, patches the logic, and sometimes recovers. Other times it gracefully walks into a wall, with the confidence of a consultant holding a laser pointer. ...

February 5, 2026 · 12 min · Zelina
Cover image

Attention Is All the Agents Need

Meetings are useful only when people listen. Anyone who has sat through a badly run management meeting knows the opposite version too: five smart people speak, nobody resolves contradictions, the loudest answer survives, and the final memo becomes a polished blend of everyone’s confusion. Congratulations. You have built an expensive consensus machine. ...

January 26, 2026 · 19 min · Zelina
Cover image

Affective Inertia: Teaching LLM Agents to Remember Who They Are

Affective Inertia: Teaching LLM Agents to Remember Who They Are A chatbot does not need to forget your name to become strange. Sometimes the stranger failure is tonal. The assistant is patient for ten turns, defensive on the eleventh, apologetic on the twelfth, and oddly cheerful on the thirteenth. Nothing in the user’s goal changed. Nothing in the product specification said “please behave like an emotionally unstable intern with excellent grammar.” Yet the agent flips. ...

January 23, 2026 · 15 min · Zelina
Cover image

Skeletons in the Proof Closet: When Lean Provers Need Hints, Not More Compute

Compute is a very convenient alibi. When an AI system fails, the modern reflex is to ask for more of it: more samples, more tokens, more search, more GPUs, more patience from whoever is paying the invoice. This habit is not always wrong. Sometimes the model really does need another attempt. Sometimes the winning answer is hiding in sample number 47. ...

January 23, 2026 · 16 min · Zelina