Same VQA Score, Different Eyes: What Fine-Grained Tests Reveal About VLMs
TL;DR for operators A vision-language model can rank well on broad question-answering benchmarks and still be materially weaker at distinguishing visually similar objects. In Understanding the Fine-Grained Knowledge Capabilities of Vision-Language Models, Ghosh, Zhang, and Schmidt show that CogVLM-Chat and LLaVA-NeXT-Vicuna-13B both score around 48% on aggregated general VQA, yet reach 77.3% and 58.1% respectively on their fine-grained evaluation—a gap of roughly 19 percentage points.1 ...