Cover image

Paging Dr. Model: When AI Runs the Workup

DxDirector-7B shows why clinical AI becomes more operationally interesting when it plans the diagnostic workflow, not merely when it answers medical questions.

August 18, 2025 · 18 min · Zelina
Cover image

Patch Tuesday for the Law: Hunting Legal Zero‑Days in AI Governance

A mechanism-first reading of Legal Zero-Days as an AI governance risk: not ordinary loophole hunting, but institutional vulnerability discovery at machine speed.

August 18, 2025 · 15 min · Zelina
Cover image

Skip or Split? How LLMs Can Make Old-School Planners Run Circles Around Complexity

A comparison-based reading of why LLM-assisted planning works best when the model predicts constrained intermediate states rather than pretending to be the planner.

August 18, 2025 · 16 min · Zelina
Cover image

Therapy, Explained: How Multi‑Agent LLMs Turn DSM‑5 Screens into Auditable Logic

DSM5AgentFlow shows that the business value of clinical LLM agents is not autonomous diagnosis, but auditable intake infrastructure.

August 18, 2025 · 17 min · Zelina
Cover image

Three’s Company: When LLMs Argue Their Way to Alpha

AlphaAgents shows how role-based LLM agents can act less like an autonomous portfolio manager and more like a structured, auditable investment committee for stock selection.

August 18, 2025 · 15 min · Zelina
Cover image

Count Us In: How Dual‑Agent LLMs Turn Math Slips into Teachable Moments

A close read of new evidence on where LLMs actually fail at math—and how dual-agent review, step labels, and tool delegation can make AI tutors more trustworthy.

August 16, 2025 · 17 min · Zelina
Cover image

Fast & Curious: How ‘Speed-First’ LLM Architectures Change the Build vs. Buy Math

A business-oriented map of efficient LLM architectures, showing which speed lever matters for latency, memory, long context, capacity, edge deployment, and agentic workloads.

August 16, 2025 · 20 min · Zelina
Cover image

Forecast: Mostly Context with a Chance of Routing

A practical reading of four LLM strategies for context-aided forecasting: diagnose failures, correct forecasts, learn from examples, and route expensive models only where they earn their keep.

August 16, 2025 · 17 min · Zelina
Cover image

Kill Switch Ethics: What the PacifAIst Benchmark Really Measures

A practical reading of PacifAIst, a benchmark that tests whether LLMs prioritise human safety when their own operational goals are on the line.

August 16, 2025 · 17 min · Zelina
Cover image

RAGulating Compliance: When Triplets Trump Chunks

A careful reading of triplet-based regulatory RAG shows why knowledge graphs matter most for auditability, strict retrieval, and navigation—not magical answer accuracy.

August 16, 2025 · 14 min · Zelina