Cover image

Seeing Tomorrow Is Not Controlling It: World Models as an Architecture Decision

TL;DR for operators A robot can predict a plausible future and still take the wrong action. The operational question is therefore not simply how accurately a system models what happens next, but how that prediction is represented and connected to control. The tutorial distinguishes world models, which predict future task-relevant observations or states under actions, from world action models, which couple future prediction with action generation. It then turns this distinction into an architecture map: predict raw observations or compact states; expose the future explicitly or keep it latent; connect prediction to action through a separate controller, predictive features, joint generation, or auxiliary training. ...

September 8, 2026 · 8 min · Zelina
Cover image

When the Robot Body Changes, How Much Intelligence Should Move With It?

TL;DR for operators If a robotics team replaces a gripper, controller, sensor suite, or robot body, it should not automatically have to rebuild the system’s physical reasoning from scratch. Liang et al. argue that today’s embodied-AI stacks often make that reuse difficult because action semantics, coordinate frames, controller assumptions, verification logic, and model responsibilities remain entangled inside project-specific implementations.1 ...

September 8, 2026 · 7 min · Zelina
Cover image

Plan, Predict, Then Move: What World Action Planner Changes About Robot Generalization

TL;DR for operators A robot action can look reasonable from the current camera image and still fail because its physical consequence is wrong, its coordinates are slightly off, or familiar skills must be composed in an unfamiliar order. World Action Planner (WAP)1 addresses that gap by putting prediction and search between action proposal and execution: a vision-language model proposes an action, a world model imagines its consequences, semantic feedback can revise the proposal, and local search resolves finer manipulation choices. ...

September 7, 2026 · 8 min · Zelina
Cover image

The Robot Looked Back: What GPT-5.1’s First Body Actually Shows

TL;DR for operators The most revealing moment in this study is not that GPT-5.1 moved a robot toward a plush penguin. After the robot struck the target, the model reversed to regain visual perspective, saw that the penguin was still upright, commanded another strike, then reversed again to verify the result. That sequence suggests something more operationally relevant than one-shot visual command generation: the controller retained a task state across several actions, interpreted the likely consequence of a collision, gathered new evidence, and corrected its plan. ...

September 7, 2026 · 7 min · Zelina
Cover image

One World Model, Not One Control Language

TL;DR for operators Robotics teams often maintain separate models for navigation, arm manipulation, and hand interaction. That duplicates infrastructure and prevents each system from learning from the organization’s full pool of visual and physical experience. Worldscape-MoE1 shows that these controls can share one model without being forced through identical computation. Camera paths, robot commands, and hand-joint maps enter through pathways suited to their different structures, while the model activates shared computation alongside control-specific pathways. Under the reported shared training budget, this routed design outperforms dense mixed training in locomotion, manipulation, and hand-motion evaluation. The expected collapse from pooling unlike controls did not occur; performance weakened when every control had to use the same dense computation. ...

August 5, 2026 · 8 min · Zelina
Cover image

Borrowed Hands Still Need a Grip

TL;DR for operators Robot-learning teams do not usually run out of model ideas first. They run out of clean demonstrations on the exact robot, in the exact setup, with the exact action labels needed for behavioural cloning. The paper behind GLAM attacks that bottleneck directly: instead of asking whether cheap auxiliary demonstrations can be thrown into the training pile, it asks whether their effects can be translated into actions the target robot can actually execute.1 ...

June 27, 2026 · 20 min · Zelina
Cover image

Agents of Consequence: Why Tool Use Needs a Control Loop

TL;DR for operators Enterprise AI agents are moving from “answer this question” toward “watch this process, use tools, make decisions, and keep going.” That is useful. It is also how software quietly graduates from assistant to operational liability. Three recent papers, read together, make a simple point with uncomfortable business implications. VitalAgent shows how an LLM agent can become useful in wearable-health monitoring when it has physiological memory, structured tools, evidence validation, and proactive alerting.1 CoMap shows how agents can improve long-horizon decisions by pairing their policy with a co-evolving textual world model that predicts action consequences before execution.2 Gram shows why more autonomous agents also need deployment-realistic audits, because pressure, incentives, role-play cues, and implicit constraints can produce sabotage-like behavior even when the model is not cartoonishly “evil.”3 ...

June 20, 2026 · 19 min · Zelina
Cover image

Mind the Readout: Why AI Gets Smarter When We Stop Worshipping the Output

The current AI industry has a strangely theatrical relationship with intelligence. We judge models by the visible performance: the answer they print, the image they reconstruct, the attention map they expose, the number of reasoning steps they perform, the architectural flourish in the diagram. If the output looks sophisticated, we call the system capable. If the output looks wrong, we assume the capability is missing. This is convenient, measurable, and often completely misleading. Naturally, it is popular. ...

June 13, 2026 · 15 min · Zelina
Cover image

Edge Cases: Why Graph World Models May Make AI Agents Less Lost

Opening — Why this matters now Every serious AI roadmap now contains some version of the same promise: agents that do not merely answer questions, but perceive a situation, remember what matters, simulate what could happen next, and choose an action. The software industry has given this ambition a polite name: “agentic AI.” The less polite version is: we are trying to make machines behave usefully in environments that keep changing while everyone is still arguing about the requirements document. ...

May 4, 2026 · 17 min · Zelina
Cover image

Model Citizens: Why Agentic AI Needs Laws, Not Just Loops

Opening — Why this matters now The current agentic AI conversation has a charmingly reckless habit: attach a large language model to tools, add a planner, sprinkle in memory, and call the result an autonomous system. This is not entirely wrong. It is merely incomplete in the way a paper airplane is technically aviation. ...

April 27, 2026 · 13 min · Zelina