Cover image

Following Instructions Is Not the Same as Knowing More

TL;DR for operators A multimodal model can become much better at obeying instructions without becoming much better at the underlying tasks those instructions govern. In the experiments examined here, one 8B vision-language model gains 10.58 percentage points on a targeted instruction-following benchmark and another gains 22.91 points. Yet their average results across broader STEM, VQA, OCR, and document-understanding tests move by only +0.33 and -0.21 points. ...

October 1, 2026 · 7 min · Zelina