Cover image

When Models Get Lost in Space: Why MLLMs Still Fail Geometry

MathSpatial shows that frontier multimodal models still struggle with clean geometric spatial reasoning, revealing a practical diagnostic gap for physical-world AI systems.

February 14, 2026 · 15 min · Zelina
Cover image

Breaking Things on Purpose: How CLI-Gym Teaches AI to Fix the Real World

A mechanism-first reading of CLI-Gym, a pipeline that turns working Dockerized repositories into scalable environment-repair tasks for stronger coding agents.

February 13, 2026 · 15 min · Zelina
Cover image

Checklist Capital: Reinforcing Agents Without Verifiable Rewards

How CM2 turns open-ended agent behavior into evidence-grounded checklist rewards, and why sparse reward assignment can be safer than denser step-level signals.

February 13, 2026 · 17 min · Zelina
Cover image

Game On, Agents: When Multimodality Meets the Godot Engine

GameDevBench shows why game development is a harsher test for AI agents than ordinary coding benchmarks: the hard part is not just writing code, but seeing, placing, animating, and verifying work inside a visual engine.

February 13, 2026 · 19 min · Zelina
Cover image

Lost in Translation: When 14% WER Hides a 44% Failure Rate

Why speech models can look reliable on benchmark metrics while still failing on the named entities that drive real-world routing, cost, and fairness.

February 13, 2026 · 15 min · Zelina
Cover image

No More ‘Trust Me, Bro’: Statistical Parsing Meets Verifiable Reasoning

A business-focused reading of how statistical parsing, typed grammar, and Logical Bayesian Networks could make enterprise AI answers more auditable without pretending LLMs have become theorem provers.

February 13, 2026 · 17 min · Zelina
Cover image

Proof Over Probabilities: Why AI Oversight Needs a Judge That Can Do Math

A mechanism-first reading of FORMALJUDGE, showing why safer AI-agent oversight may depend less on stronger judges and more on formally checkable constraints.

February 13, 2026 · 17 min · Zelina
Cover image

See, Plan, Snap: Why AI Can Think in Blocks but Can’t Drop Them

ScratchWorld shows that today’s multimodal GUI agents can often reason about visual programs, but still fail where business automation actually hurts: precise, reliable execution.

February 13, 2026 · 16 min · Zelina
Cover image

Think Like a Scientist: When LLMs Stop Guessing and Start Reasoning

How KeplerAgent turns LLMs from equation guessers into tool-orchestrating scientific reasoning systems—and what that means for interpretable AI in R&D.

February 13, 2026 · 15 min · Zelina
Cover image

Thinking About Thinking: When LLMs Start Writing Their Own Report Cards

RLCER shows how self-evolving rubrics can turn reinforcement learning from answer checking into process-level reasoning supervision.

February 13, 2026 · 18 min · Zelina