Cover image

Trex Marks the Spot: When AI Starts Training AI

A mechanism-first reading of TREX, an agent system that treats LLM fine-tuning as an iterative research workflow rather than a glorified hyperparameter search.

April 16, 2026 · 16 min · Zelina
Cover image

When Maps Start Thinking: GeoAgentBench and the Audit of Spatial AI

GeoAgentBench shows why serious spatial AI must be tested by execution, parameter discipline, and final map verification—not by how convincingly an agent describes a workflow.

April 16, 2026 · 17 min · Zelina
Cover image

Benchmarking the Benchmarks: When AI Safety Metrics Stop Meaning Anything

A sharper reading of AISafetyBenchExplorer, showing why AI safety evaluation now suffers less from benchmark scarcity than from metric drift, stale infrastructure, and weak benchmark governance.

April 15, 2026 · 16 min · Zelina
Cover image

Evolve or Die Trying: When LLMs Stop Writing Code and Start Designing Algorithms

BEAM shows that useful LLM algorithm design is less about clever prompting and more about structured search, reusable memory, and evaluation that actually resembles solver construction.

April 15, 2026 · 18 min · Zelina
Cover image

From Words to Workflows: Why AI Still Struggles to Think Like an Operations Research Analyst

A close reading of Text2Model shows why LLMs can draft optimization models, but still need validation layers before they can be trusted in business decision workflows.

April 15, 2026 · 15 min · Zelina
Cover image

Learning on Autopilot? Not Quite — How PAL Turns Passive Videos into Active Intelligence

A mechanism-first reading of PAL, an AI learning platform that turns lecture videos into adaptive questioning, learner-state tracking, and personalized post-lesson reinforcement.

April 15, 2026 · 14 min · Zelina
Cover image

Routing Without Running Out: How Bilevel Optimization Rewires EV Logistics

A mechanism-first reading of how bilevel optimization makes electric vehicle routing more scalable by using routing cost as a cheap but imperfect guide.

April 15, 2026 · 15 min · Zelina
Cover image

The Memory Isn’t Broken — It’s Flat: Why LLMs Need to ‘Draw’ to Remember

A mechanism-first reading of dual-trace memory encoding and why enterprise AI agents may need richer contextual traces, not just larger memory stores.

April 15, 2026 · 15 min · Zelina
Cover image

The Search That Remembers: Training AI Without Answers

How Cycle-Consistent Search turns the search trajectory itself into a reward signal for training AI agents when gold answers are unavailable.

April 15, 2026 · 17 min · Zelina
Cover image

Epistemic Infrastructure: Why Your AI Knows Less Than It Thinks

A measured reading of OIDA: why organizational AI needs memory that tracks decisions, contradictions, and open questions—not just better retrieval.

April 14, 2026 · 15 min · Zelina