When Models Get Lost in Space: Why MLLMs Still Fail Geometry
MathSpatial shows that frontier multimodal models still struggle with clean geometric spatial reasoning, revealing a practical diagnostic gap for physical-world AI systems.
MathSpatial shows that frontier multimodal models still struggle with clean geometric spatial reasoning, revealing a practical diagnostic gap for physical-world AI systems.
A mechanism-first reading of CLI-Gym, a pipeline that turns working Dockerized repositories into scalable environment-repair tasks for stronger coding agents.
How CM2 turns open-ended agent behavior into evidence-grounded checklist rewards, and why sparse reward assignment can be safer than denser step-level signals.
GameDevBench shows why game development is a harsher test for AI agents than ordinary coding benchmarks: the hard part is not just writing code, but seeing, placing, animating, and verifying work inside a visual engine.
Why speech models can look reliable on benchmark metrics while still failing on the named entities that drive real-world routing, cost, and fairness.
A business-focused reading of how statistical parsing, typed grammar, and Logical Bayesian Networks could make enterprise AI answers more auditable without pretending LLMs have become theorem provers.
A mechanism-first reading of FORMALJUDGE, showing why safer AI-agent oversight may depend less on stronger judges and more on formally checkable constraints.
ScratchWorld shows that today’s multimodal GUI agents can often reason about visual programs, but still fail where business automation actually hurts: precise, reliable execution.
How KeplerAgent turns LLMs from equation guessers into tool-orchestrating scientific reasoning systems—and what that means for interpretable AI in R&D.
RLCER shows how self-evolving rubrics can turn reinforcement learning from answer checking into process-level reasoning supervision.