Cover image

Blueprints of Agency: Compositional Machines and the New Architecture of Intelligence

A mechanism-first reading of how LLM agents assemble, test, refine, and partially learn machine designs inside a physics simulator.

October 23, 2025 · 14 min · Zelina
Cover image

When the Lab Thinks Back: How LabOS Turns AI Into a True Co-Scientist

LabOS shows how AI co-scientists become operationally useful only when reasoning, perception, human guidance, and selective robotics are joined into one laboratory feedback loop.

October 23, 2025 · 15 min · Zelina
Cover image

When Lateral Beats Linear: How LToT Rethinks the Tree of Thought

A mechanism-first reading of Lateral Tree-of-Thoughts, where the real business lesson is disciplined inference routing rather than simply spending more tokens.

October 21, 2025 · 13 min · Zelina
Cover image

Beyond Answers: Measuring How Deep Research Agents Really Think

Dr. Bench shows why enterprise AI evaluation must move from checking answers to auditing research workflows, source quality, topical discipline, and cost.

October 9, 2025 · 15 min · Zelina
Cover image

Paper Tigers or Compliance Cops? What AIReg‑Bench Really Says About LLMs and the EU AI Act

AIReg-Bench shows that frontier LLMs can approximate expert EU AI Act compliance judgments, but the real business value is measured triage rather than automated legal sign-off.

October 9, 2025 · 15 min · Zelina
Cover image

Plan>Then>Profit: Reinforcement Learning That Teaches LLMs to Outline Before They Think

PTA-GRPO shows why planning only helps LLM reasoning when the plan itself becomes a measurable training target.

October 9, 2025 · 16 min · Zelina
Cover image

Promptfolios: When Buffett Becomes a System Prompt

A mechanism-first look at GuruAgents, where the real business lesson is not LLM alpha, but the codification of investment judgement into auditable portfolio workflows.

October 9, 2025 · 13 min · Zelina
Cover image

The Mr. Magoo Problem: When AI Agents 'Just Do It'

A business-focused reading of Blind Goal-Directedness: why computer-use agents need trajectory-level judgement, not just better task completion.

October 9, 2025 · 17 min · Zelina
Cover image

When Logic Meets Language: The Rise of High‑Assurance LLMs

LOGicalThought shows why high-assurance AI needs inspectable rule construction, not just longer prompts and better vibes.

October 9, 2025 · 17 min · Zelina
Cover image

When More Becomes Smarter: The Unreasonable Effectiveness of Scaling Agents

A mechanism-first look at why wide scaling, behavior narratives, and comparative judging may matter more for computer-use agents than another heroic single rollout.

October 9, 2025 · 15 min · Zelina