TL;DR for operators

A robot-learning pipeline may use a vision-language model to decide which demonstrations to retain, whether a policy earned a reward, whether an execution should be retried, or which policy performed better. That makes the reliability of the judge part of the system, not a reporting detail.

Because success and failure frequencies can differ sharply across datasets, ordinary accuracy can reward a model that mostly predicts the majority class. FailBench therefore emphasizes balanced accuracy. Its strongest tested detector reaches only 0.77 macro balanced accuracy across the two-class subsets. More unexpectedly, every purpose-built failure detector performs below nearly all of the general-purpose VLMs, and all five specialists with matched comparisons score below their own base models.

The operational lesson is not that specialization is inherently harmful. It is that specialization should be treated as a hypothesis that must survive distribution shift. Teams deploying VLM judges should test them across independently collected tasks, cameras, robots, and failure patterns, compare fine-tuned models against their bases, and preserve human review where mistakes can contaminate rewards or training data.

The paper also identifies a narrower deployment lever. When decisive evidence occupies a small region of the image, cropping to that region raises the strongest detector from 0.773 to 0.797 macro balanced accuracy without retraining. The improvement is statistically supported but uneven across sources, and it does not remove the deeper weakness on contact-dependent outcomes.

A robot pipeline can fail even when the robot did not

Suppose a manipulation policy finishes an attempt and a VLM decides whether the task succeeded. That decision can then become a label, a reward, a filtering rule, or a trigger for another action.

The attractive assumption is that this judgment is comparatively easy: the robot has already acted, the video exists, and the model only has to inspect the outcome. But if the judge is wrong after the robot, camera, object arrangement, or failure pattern changes, the error propagates downstream. A successful attempt may be discarded. A failed trajectory may be retained as training data. A policy may receive the wrong reward or skip a necessary recovery.

FailBench, introduced by Navasardyan, Danielyan, and Davtyan,1 tests this problem across 2,197 completed manipulation attempts from 14 public sources. Twelve sources are real-world and two are simulated, with 1,176 failures and 1,021 successes. Thirteen VLM-based detectors are evaluated under one common protocol.

The strongest result is therefore also the first warning: Gemini 3 Flash reaches 0.77 macro balanced accuracy. At benchmark scale, even the leading tested judge still gets roughly one quarter of decisions wrong.

Specialization does not survive the cross-source test automatically

A reasonable procurement assumption is that a detector fine-tuned specifically for robot failures should be safer than its general-purpose parent model.

FailBench finds the opposite pattern under its cross-source protocol. Every purpose-built detector scores below every general-purpose VLM except the smallest one. More tellingly, the paper compares five specialists directly with their own base models while holding prompt, decoding, and input recipe fixed.

Specialist Base model Specialist Base Delta
Guardian (thinking) InternVL3-8B 0.627 0.629 -0.002
RoboReward-8B Qwen3-VL-8B-Instruct 0.619 0.675 -0.057
ViFailback-8B Qwen3-VL-8B-Instruct 0.587 0.675 -0.088
RoboFAC-7B Qwen2.5-VL-7B-Instruct 0.514 0.555 -0.040
FailSense-Calvin-3B PaliGemma2-3B 0.503 0.506 -0.003

These comparisons isolate the fine-tuning difference more cleanly than the overall leaderboard. They do not establish that robotics fine-tuning generally reduces capability. They show that these five failure-specialized models did not preserve an advantage when evaluated across independently collected sources.

For operators, that changes the validation question. “Was this model trained for robot failures?” is weaker evidence than “Does the specialization still beat the base model on tasks and environments outside the development distribution?”

Metric choice can hide a detector that barely discriminates

Cross-source testing is only useful if the score itself measures the intended behavior.

FailBench highlights how plain accuracy can mislead on failure-heavy datasets. On the published RoboFAC split, a detector that always predicts failure would already achieve 0.797 ordinary accuracy, close to the reported 0.806. On the cited ViFailback split, the analogous all-failure baseline reaches 0.890.

A high raw accuracy number can therefore coexist with very weak discrimination between success and failure.

Balanced accuracy gives the two classes equal weight and is consequently the more informative metric for this benchmark. For model procurement or internal evaluation, the broader operational point is straightforward: report class composition, inspect trivial majority-class baselines, and use metrics that remain meaningful when outcome frequencies shift.

Contact state, not just visual recognition, is the hard case

The errors are not distributed uniformly across tasks.

REASSEMBLE, a contact-rich assembly source, is the hardest real-world subset. Average detector performance is about 0.52 balanced accuracy, and no evaluated model exceeds 0.60. By contrast, tasks whose outcomes can be inferred from coarse object identity or visible motion are generally easier.

The paper interprets this as an evidence problem. Determining that an object moved to the correct region can require relatively coarse visual information. Determining that two parts are properly inserted, clipped, seated, or physically engaged can depend on a small contact region and subtle state differences.

That mechanism is plausible but should not be overstated. The detailed manual diagnosis is non-statistical and concentrates mainly on Gemma-4-31B-it traces and failed episodes. FailBench does not contain a benchmark-wide taxonomy proving that contact reasoning causes the full performance gap.

Still, the operational distinction is useful. A team evaluating bin placement and a team evaluating precision insertion are not asking the visual judge to recover the same kind of evidence.

Cropping the evidence helps, but only modestly

If the decisive state occupies a small portion of a wide scene, one response is to change the input rather than retrain the detector.

The paper tests a localization pipeline in which a VLM first identifies the outcome-relevant region from the instruction and several frames. Fixed-camera inputs are then cropped to that region and passed through the detector with otherwise unchanged settings.

For Gemini 3 Flash, macro balanced accuracy rises from 0.773 to 0.797. Across the paired evaluation, cropping fixes 223 previously wrong decisions and breaks 160 previously correct ones; the difference is significant under an exact McNemar test with $p=0.0015$.

This is main evidence for an input-level intervention, not evidence that localization solves robot failure detection. Gains vary substantially by source, and several subsets worsen after cropping. REASSEMBLE improves only from 0.567 to 0.591.

The distinction matters. Localization can remove irrelevant context and enlarge the region containing decisive evidence. It cannot create information that the visual stream does not encode clearly enough, nor can it supply the physical-contact inference the model lacks.

What Cognaptus would change in deployment

The paper directly supports three changes to evaluation practice.

First, validate the judge across independently collected sources before allowing its outputs to become rewards, labels, filters, or recovery signals. A home-benchmark result is weak evidence for performance after robot, task, camera, or failure-generation conditions change.

Second, regression-test specialized models against their own bases. Fine-tuning should earn its deployment slot by improving the shifted evaluation set, not merely by having a domain-specific training objective.

Third, treat input design as part of model performance. Evidence localization is inexpensive relative to another training cycle and can help when the relevant object occupies a small fraction of the scene.

Cognaptus would add a stronger operational boundary: for contact-heavy workflows, a vision-only VLM should not automatically be treated as the final arbiter. Force-torque signals, audio, dedicated physical-state models, or human review may carry information unavailable in the image. FailBench does not test whether these modalities close the gap, so that remains a deployment hypothesis rather than a result of the paper.

The benchmark measures post-hoc success, not robot reliability as a whole

FailBench covers binary judgments after execution. It does not evaluate planning errors, online anticipation of failure, dense progress estimation, or explanations of failure type.

Its embodiment coverage is also narrow: tabletop parallel-jaw arms dominate the benchmark, with limited bimanual data and no humanoid, mobile-manipulation, or dexterous-hand evaluation. Four source datasets contain force-torque or audio signals, but the benchmark uses visual input plus the instruction.

Those limits matter most when transferring the findings to industrial settings where success depends on force, compliance, tactile engagement, or richer embodiments.

Within its scope, however, the result is clear. A VLM judge can look strong on familiar data while remaining unreliable across independently collected environments, and a domain-specific fine-tune can perform worse than the model it started from.

The deployment question is therefore not whether a system has a robot failure detector. It is whether that detector has been tested on the variation its decisions will actually govern.

Cognaptus: Automate the Present, Incubate the Future.


  1. Zaruhi Navasardyan and Tatul Danielyan and Hrant Davtyan (2026). FailBench: How Reliable are VLMs at Judging Robot Task Success?. arXiv:2609.03611. https://arxiv.org/abs/2609.03611 ↩︎