Cover image

The Goats in the Machine: Why AI Agents Need Contracts, Not Personalities

A practical reading of two new agent papers showing why enterprise AI should be judged by observable behaviour and runtime contracts, not human-like performance theatre.

June 16, 2026 · 15 min · Zelina
Cover image

Bond Before Brain: What Actually Drives Molecular MPNNs

A mechanism-first reading of a factorial molecular GNN benchmark showing why message construction deserves more attention than architectural nameplates.

June 15, 2026 · 16 min · Zelina
Cover image

Cheap Seats, Sharp Eyes: Reward-Hack Detection Without the Frontier Judge

A small trajectory encoder nearly matches a frontier LLM judge on reward-hack detection, but only when it can read the reasoning-rich trace.

June 15, 2026 · 6 min · Zelina
Cover image

Judge, Jury, and Calibration: Why AI Evaluation Needs Anchors

A practical reading of two arXiv papers showing why AI works best in high-stakes evaluation when it is anchored to human evidence and audited for real engagement.

June 15, 2026 · 14 min · Zelina
Cover image

Pre-Review, Not Peer Review: The Drafting Gate AI Actually Earns

A mechanism-first reading of AI-Paper-Review shows why AI review is useful as a pre-submission quality gate, not as a substitute for human peer review.

June 15, 2026 · 18 min · Zelina
Cover image

Split Before You Scale: Why Useful AI Starts by Sorting the Mess

A business-focused reading of three arXiv papers showing why scalable AI depends on decomposing structure, uncertainty, and supervision before optimisation.

June 15, 2026 · 16 min · Zelina
Cover image

Statecraft, Not Scorecards: Why Reliable AI Lives on the Path

A practical reading of two arXiv papers showing why AI reliability depends on the states models visit and the trajectories evaluators inspect.

June 15, 2026 · 3 min · Zelina
Cover image

The Chatbot Passed the Test. Then It Bowed Too Low.

NICE shows why aggregate social-intelligence scores can hide the communication failures that matter most in real deployments.

June 15, 2026 · 18 min · Zelina
Cover image

Feedback, Not Freefall: Why LLM Writing Tools Need a Teacher in the Loop

A large K-12 writing study shows that LLM feedback works best as a teacher-mediated workflow, not as a replacement chatbot with better grammar.

June 14, 2026 · 17 min · Zelina
Cover image

Frame Before You Aim: Why AI Needs the Right Reference Point

A cross-domain look at why useful AI systems need adaptation layers that translate models, protocols, and rankings into the realities they are meant to serve.

June 14, 2026 · 15 min · Zelina