Cover image

Who Checks the Robot Judge? Cross-Source Reliability in VLM Failure Detection

TL;DR for operators A robot-learning pipeline may use a vision-language model to decide which demonstrations to retain, whether a policy earned a reward, whether an execution should be retried, or which policy performed better. That makes the reliability of the judge part of the system, not a reporting detail. Because success and failure frequencies can differ sharply across datasets, ordinary accuracy can reward a model that mostly predicts the majority class. FailBench therefore emphasizes balanced accuracy. Its strongest tested detector reaches only 0.77 macro balanced accuracy across the two-class subsets. More unexpectedly, every purpose-built failure detector performs below nearly all of the general-purpose VLMs, and all five specialists with matched comparisons score below their own base models. ...

October 4, 2026 · 7 min · Zelina
Cover image

Human Demonstrations Need a Relevance Filter Before VLA Post-Training

TL;DR for operators A robotics team may have a small, expensive set of demonstrations from its target robot and a much larger pool of cheaper human demonstrations of apparently similar tasks. The tempting move is to combine them. The paper’s randomized simulation results show why that decision needs more control: robot-only training averages 0.34 task success, while randomly mixed human data falls to 0.32. Selecting human demonstrations by relevance raises the average to 0.40, and adding sample-specific weighting raises it further to 0.42. ...

September 25, 2026 · 7 min · Zelina
Cover image

When the Robot Body Changes, How Much Intelligence Should Move With It?

TL;DR for operators If a robotics team replaces a gripper, controller, sensor suite, or robot body, it should not automatically have to rebuild the system’s physical reasoning from scratch. Liang et al. argue that today’s embodied-AI stacks often make that reuse difficult because action semantics, coordinate frames, controller assumptions, verification logic, and model responsibilities remain entangled inside project-specific implementations.1 ...

September 8, 2026 · 7 min · Zelina
Cover image

Take Over, Then Let Go: AutoIntervene Makes Robot Recovery a Training Signal

TL;DR for operators When an imitation-learned robot drifts outside the situations represented in its demonstrations, detecting that something looks unusual solves only half the operational problem. A deployment system must also decide when human takeover is warranted, when the recovery has progressed far enough to return control, and whether that intervention can improve the next policy rather than disappear as one-off operational labor. ...

August 22, 2026 · 8 min · Zelina