Cover image

Assert Less, Observe More: AICL and the New QA Stack for LLM Apps

A practical reading of AICL as a testability layer for LLM applications, where QA shifts from exact assertions to observable, replayable behaviour.

August 31, 2025 · 17 min · Zelina
Cover image

From Chat Logs to Goal Logs: OnGoal’s Playbook for Goal‑Truthful LLMs

OnGoal shows why long LLM conversations need goal observability, not just longer context windows.

August 31, 2025 · 16 min · Zelina
Cover image

Prolog & Paycheck: When Tax AI Shows Its Work

A mechanism-first reading of why tax AI becomes more useful when LLMs translate rules into executable logic, defer when uncertain, and price mistakes like real liabilities.

August 31, 2025 · 15 min · Zelina
Cover image

Rollouts, Not GPUs: Why AWorld’s 14.6× Speedup Rewires Agent Training

AWorld shows that the practical bottleneck in agent training is not only model capacity or gradient compute, but scalable experience generation.

August 31, 2025 · 15 min · Zelina
Cover image

Vitals, Not Vibes: Inside the New Anatomy of Personal Health Agents

A mechanism-first look at why personal health AI needs specialised agents, grounded tools, orchestration, memory, and evaluation before it can move beyond wellness chatbot theatre.

August 31, 2025 · 15 min · Zelina
Cover image

Benchmarks with Benefits: What DeepScholar-Bench Really Measures

DeepScholar-Bench shows that research agents should be judged by coverage, source quality, and citation support—not by how convincingly they format a report.

August 30, 2025 · 14 min · Zelina
Cover image

Edge of Reason: Orchestrating LLMs Without a Conductor

A mechanism-first look at Symphony, a decentralised multi-agent LLM framework that routes reasoning work through capability matching instead of a central conductor.

August 30, 2025 · 16 min · Zelina
Cover image

Faking It to Make It: When Synthetic Data Actually Works

A practical map for deciding when generative synthetic data is useful, when it is theatre, and what must be evaluated before it touches production.

August 30, 2025 · 18 min · Zelina
Cover image

MoE Money, MoE Problems? FinCast Bets Big on Foundation Models for Markets

FinCast shows how a finance-specific time-series foundation model can improve forecasting accuracy, but not yet prove tradable alpha.

August 30, 2025 · 16 min · Zelina
Cover image

Who Watches the Watchers? Weak-to-Strong Monitoring that Actually Works

A practical reading of monitor red teaming for LLM agents, showing why scaffolding and escalation policy matter more than simply giving monitors more context.

August 30, 2025 · 17 min · Zelina