Cover image

Plan, Predict, Then Move: What World Action Planner Changes About Robot Generalization

TL;DR for operators A robot action can look reasonable from the current camera image and still fail because its physical consequence is wrong, its coordinates are slightly off, or familiar skills must be composed in an unfamiliar order. World Action Planner (WAP)1 addresses that gap by putting prediction and search between action proposal and execution: a vision-language model proposes an action, a world model imagines its consequences, semantic feedback can revise the proposal, and local search resolves finer manipulation choices. ...

September 7, 2026 · 8 min · Zelina
Cover image

Take Over, Then Let Go: AutoIntervene Makes Robot Recovery a Training Signal

TL;DR for operators When an imitation-learned robot drifts outside the situations represented in its demonstrations, detecting that something looks unusual solves only half the operational problem. A deployment system must also decide when human takeover is warranted, when the recovery has progressed far enough to return control, and whether that intervention can improve the next policy rather than disappear as one-off operational labor. ...

August 22, 2026 · 8 min · Zelina
Cover image

Borrowed Hands Still Need a Grip

TL;DR for operators Robot-learning teams do not usually run out of model ideas first. They run out of clean demonstrations on the exact robot, in the exact setup, with the exact action labels needed for behavioural cloning. The paper behind GLAM attacks that bottleneck directly: instead of asking whether cheap auxiliary demonstrations can be thrown into the training pile, it asks whether their effects can be translated into actions the target robot can actually execute.1 ...

June 27, 2026 · 20 min · Zelina
Cover image

Fold Me Once: When the Demonstration Becomes the Robot Interface

TL;DR for operators Instant-Fold is not mainly a “robot folds shirts” paper. That is the demo-friendly surface layer, and robotics papers do need a surface layer. The more useful idea is that a single demonstration can work as an operational interface for deformable tasks where language is too thin, checklists are too brittle, and final-state labels hide the important part: how the object got there.1 ...

June 25, 2026 · 18 min · Zelina
Cover image

Context Is Not a Costume: Why Strong Agents Still Fail on Contact

The agent looks ready. Then reality answers back. The current AI-agent story is conveniently simple. Take a powerful foundation model, wrap it in tools, give it a workflow, add a polite system prompt, and call the result “ready for deployment.” Reality, as usual, has poor manners. Two recent arXiv papers examine very different agent settings. One studies whether multimodal AI agents can align their behavior with the cognitive age of child users. The other studies whether behavior foundation models for imitation learning can remain robust when the physical dynamics of an environment shift after training. They do not share a benchmark, a model class, or even the same deployment domain. That is precisely why they are useful together. ...

May 29, 2026 · 14 min · Zelina
Cover image

ImplicitRDP: When Robots Stop Guessing and Start Feeling

Robots are very good at looking confident. Put a camera on a robot arm, train it with enough demonstrations, and it may glide toward a box, a switch, or a tool with the calm precision of something that understands the world. Then contact happens. The fingertip presses too hard. The switch has not actually toggled. The object slips, bends, jams, or quietly enters the expensive category known as “damaged inventory.” ...

December 13, 2025 · 17 min · Zelina
Cover image

Learning by X-ray: When Surgical Robots Teach Themselves to See in Shadows

X-rays are useful because they are cheap, familiar, and already sitting in the operating room. They are also, inconveniently, shadows. That is the central tension in Investigating Robot Control Policy Learning for Autonomous X-ray-guided Spine Procedures, a paper that asks whether a robot policy can plan vertebroplasty cannula trajectories from only bi-planar X-ray views—one anterior-posterior view, one lateral view—without CT-based navigation, registration, or a lovingly over-engineered suite of intra-operative infrastructure.1 ...

November 9, 2025 · 14 min · Zelina