Cover image

Who Checks the Robot Judge? Cross-Source Reliability in VLM Failure Detection

TL;DR for operators A robot-learning pipeline may use a vision-language model to decide which demonstrations to retain, whether a policy earned a reward, whether an execution should be retried, or which policy performed better. That makes the reliability of the judge part of the system, not a reporting detail. Because success and failure frequencies can differ sharply across datasets, ordinary accuracy can reward a model that mostly predicts the majority class. FailBench therefore emphasizes balanced accuracy. Its strongest tested detector reaches only 0.77 macro balanced accuracy across the two-class subsets. More unexpectedly, every purpose-built failure detector performs below nearly all of the general-purpose VLMs, and all five specialists with matched comparisons score below their own base models. ...

October 4, 2026 · 7 min · Zelina