Same Answer, Different Risk: Visual Semantic Entropy for VLM Review Routing
TL;DR for operators A visual assistant gives the same confident answer several times. That consistency may seem sufficient for automatic acceptance, but it shows only that repeated decoding produced the same response—not that the underlying visual interpretation is stable. Variability can also come from the wrong place. When both the image and question wording are changed, paraphrase choice may drive the resulting answer clusters more strongly than the visual changes. In the paper’s joint-perturbation analysis, text purity exceeds image purity for every reported model and split. A high uncertainty score may therefore indicate prompt sensitivity rather than visual ambiguity. ...