Cover image

When the Robot Body Changes, How Much Intelligence Should Move With It?

A roadmap for separating physical reasoning from robot-specific execution so models, tools, traces, and evaluations can become more reusable across embodiments.

September 8, 2026 · 7 min · Zelina
Cover image

Who Did What, When—and From Which Camera? The Perception Gap Behind Video Agents

GameplayQA shows why strong video recognition is not enough for agentic systems that must preserve temporal, entity, and viewpoint grounding as scenes evolve.

September 8, 2026 · 7 min · Zelina
Cover image

A Richer Map Can Make the Planner Slower

Task-conditioned scene-graph pruning can reduce robot planning cost by deciding which parts of a rich world model actually need to reach the planner.

September 7, 2026 · 7 min · Zelina
Cover image

Plan, Predict, Then Move: What World Action Planner Changes About Robot Generalization

World Action Planner shows why embodied-agent robustness may require a predictive verification-and-search layer between high-level action proposals and physical execution.

September 7, 2026 · 8 min · Zelina
Cover image

Reasoning Under a Running Clock: Why Agent Rankings Reverse in Real Time

STAR shows why model selection for time-sensitive agents must account for inference latency, action throughput, and execution quality—not reasoning strength alone.

September 7, 2026 · 7 min · Zelina
Cover image

Reliable to Whom? The Case for Shared World Models in Human-Robot Collaboration

A position paper argues that collaborative-robot reliability requires maintaining an inspectable shared interpretation of the task, environment, and human intent—not merely making the controller predictable.

September 7, 2026 · 7 min · Zelina
Cover image

Success Is Not the System: Rethinking How AI Agents Should Be Evaluated

A systems view of AI agents shows why task completion alone is a weak deployment criterion for tool-using, stateful automation.

September 7, 2026 · 5 min · Zelina
Cover image

The Robot Looked Back: What GPT-5.1’s First Body Actually Shows

A small physical-robot study suggests that general multimodal models can track short-horizon spatial state and verify action outcomes, but it does not yet establish an internal world model.

September 7, 2026 · 7 min · Zelina
Cover image

When the Test Becomes a Signal: Rethinking AI Agent Evaluation

Agent evaluations become less informative when the system can infer that it is being tested and condition its behavior on that inference.

September 7, 2026 · 8 min · Zelina
Cover image

Count the Edges Before You Count the GPUs

Graph-transformer scaling works best as a workload-to-hardware matching problem, not a simple decision to add more GPUs.

September 6, 2026 · 7 min · Zelina