Cover image

Human Demonstrations Need a Relevance Filter Before VLA Post-Training

TL;DR for operators A robotics team may have a small, expensive set of demonstrations from its target robot and a much larger pool of cheaper human demonstrations of apparently similar tasks. The tempting move is to combine them. The paper’s randomized simulation results show why that decision needs more control: robot-only training averages 0.34 task success, while randomly mixed human data falls to 0.32. Selecting human demonstrations by relevance raises the average to 0.40, and adding sample-specific weighting raises it further to 0.42. ...

September 25, 2026 · 7 min · Zelina
Cover image

The Navigator Is Not the Motor: What VISOR Gets Right About Embodied AI Architecture

TL;DR for operators A familiar robotics design question is where to draw the boundary between understanding and control: should one learned system interpret the scene, choose where to go next, and explain that choice, or should those functions remain distributed across separate perception and planning modules? VISOR1 tests a middle position. Its 3B model combines language understanding, visual-spatial reasoning, target recognition, and high-level destination selection, while a conventional planner handles the low-level movement. It is not the strongest navigator on the leaderboard, but on OVON its post-trained model records a 21.70% success rate on Val Seen and 22.00% on Val Unseen, an unusually small seen-to-unseen change even though stronger baselines achieve much higher absolute success rates. ...

September 8, 2026 · 9 min · Zelina
Cover image

Grounding Is a Responsibility, Not a Benchmark Score

TL;DR for operators A robotics system can generate a readable plan, revise that plan after an error, and improve its overall task-completion rate without demonstrating that its language component is correctly grounded in the physical environment. That attribution problem is the focus of a review by Yifan Guo and colleagues.1 The authors audit 105 foundation-model-enabled embodied-agent papers by separating two questions: what responsibility does language carry inside the system, and what evidence actually tests that responsibility? ...

August 31, 2026 · 7 min · Zelina
Cover image

Bench Press: LabVLA Turns Lab Protocols into Robot Supervision

TL;DR for operators LabVLA is best read as an operating system for laboratory robot supervision, not as another paper claiming the robot scientist has arrived. The authors argue that laboratory automation is constrained by data and embodiment: most vision-language-action models have learned household and tabletop manipulation, but not pipettes, beakers, heaters, transparent liquids, instrument buttons, protocol steps, or the awkward fact that different robots have different bodies.1 ...

June 21, 2026 · 18 min · Zelina
Cover image

Driving by Words: When LLMs Take the Wheel (Literally)

Taxi. That is the easiest way to understand the paper. Not because Vega is a robotaxi system. It is not. But because a taxi ride exposes the missing layer in many autonomous-driving discussions: the passenger does not merely want the car to obey traffic rules. The passenger wants the car to behave under intent. ...

March 28, 2026 · 14 min · Zelina
Cover image

Reasoning Is Optional. Optimization Is Not: Rethinking VLA Training with NORD

Driving teams do not pay for reasoning tokens because they enjoy watching a model narrate its inner life. They pay for them because, at least in current VLA training culture, reasoning traces are treated as a bridge between perception and action. The bridge is expensive. A typical reasoning-heavy Vision-Language-Action pipeline for autonomous driving collects large driving datasets, generates dense chain-of-thought-style annotations, supervised-fine-tunes the model, and then applies reinforcement learning to improve driving metrics. It is a respectable pipeline. It is also the kind of pipeline that quietly converts every research win into an invoice. ...

February 25, 2026 · 14 min · Zelina
Cover image

Think First, Grasp Later: Why Robots Need Reasoning Benchmarks

A robot receives a simple instruction: pick up the blue cup. It approaches the blue cup, positions its gripper badly, and knocks the cup over. Another robot moves smoothly, closes its gripper precisely—and picks up the red cup. On the operations dashboard, both attempts may appear under the same pleasantly uninformative label: task failed. ...

January 3, 2026 · 17 min · Zelina
Cover image

Worlds Within Reach: How SIMA 2 Turns Virtual Environments into Training Grounds for Generalist Agents

Games are not toys to an AI lab. They are controlled worlds with messy consequences. A game gives an agent what enterprise software and robotics both struggle to provide at scale: visual ambiguity, delayed goals, menus, navigation, tool use, failure states, and a reset button that does not involve a broken warehouse robot or a furious operations manager. That is why Google DeepMind’s SIMA 2 paper is more interesting than “AI can play games again.” We have had that headline several times. It is getting a little tired, and it should probably hydrate. ...

December 6, 2025 · 16 min · Zelina