TL;DR for operators

A multimodal model can become much better at obeying instructions without becoming much better at the underlying tasks those instructions govern.

In the experiments examined here, one 8B vision-language model gains 10.58 percentage points on a targeted instruction-following benchmark and another gains 22.91 points. Yet their average results across broader STEM, VQA, OCR, and document-understanding tests move by only +0.33 and -0.21 points.

The operational consequence is straightforward: treat instruction compliance and task capability as separate evaluation layers. Deterministic checks for formatting, language, grounding, numerical presentation, and instruction conflicts can make regression testing and reinforcement learning more reproducible. But a stronger compliance score should not replace independent capability testing.

The paper also turns visible instructions inside images into an explicit reliability test. That matters for assistants that read screenshots, documents, charts, interfaces, or other images containing text: the image can contain something that looks like an instruction even when the user asked the model to do something else.

A model can obey better without knowing more

Suppose a product team fine-tunes a vision-language assistant because users need outputs in the correct language, valid JSON, a required numerical format, or a particular grounding representation. After training, the compliance score rises sharply.

What exactly improved?

That question matters because deployment decisions often combine two different properties: whether the model can perform a task and whether it follows the rules surrounding that task. A model that already recognizes the correct object may still fail because it emits an invalid bounding-box format. A model that understands a chart may still ignore a requested answer wrapper. Conversely, teaching those behaviors more reliably does not necessarily add much new visual reasoning capability.

MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models1 gives unusually clean evidence for separating the two.

After training with the paper’s instruction-hijacking examples, Qwen3-VL-8B-Instruct rises from 78.73 to 89.31 on MM-IFEval-Pro. InternVL3.5-8B rises from 64.64 to 87.55.

Across nine broader multimodal benchmarks, however, the averages barely move:

Model MM-IFEval-Pro before After Broader benchmark average before After
Qwen3-VL-8B-Instruct 78.73 89.31 66.83 67.16
InternVL3.5-8B 64.64 87.55 63.65 63.44

The strongest evidence is therefore about better instruction following, not a large increase in general multimodal intelligence.

That distinction should affect how post-training evaluations are reported internally. A compliance benchmark can justify claims about compliance under its tested conditions. It should not automatically justify a broader claim that the underlying model has become substantially more capable.

Instruction following becomes a set of executable tests

The benchmark’s more interesting design decision is not simply adding harder prompts. It decomposes instruction following into requirements that software can check.

A rule-verifiable constraint is an instruction requirement whose satisfaction can be tested by an executable function rather than primarily scored by another language model. Examples include paragraph counts, prohibited substrings, valid JSON, decimal precision, language requirements, bounding-box formatting, and numerical answer presentation.

MM-IFEval-Pro also makes those checks language-aware. Chinese and English do not share identical notions of words, characters, or sentence segmentation, so the benchmark implements language-specific verification rather than translating a prompt and reusing an English-oriented rule.

That gives the benchmark two roles.

First, it is an evaluation system. Accuracy is the mean fraction of attached constraints that a response satisfies.

Second, the same checks become training signals. The authors build 13,547 training samples, averaging three constraints per sample, and use GRPO reinforcement learning. Each rollout receives reward according to the proportion of its constraints that pass.

The training evidence supports that design beyond the benchmark itself. For Qwen3-VL-8B-Instruct, MM-IFEval, MIA, and IFEval all improve after training. InternVL3.5-8B shows the same direction on those external instruction-following evaluations.

This is where deterministic verification has direct operational value. A requirement such as “return valid JSON” does not need a second model to decide whether the output complied. Neither does a precise bounding-box schema. When a product requirement can be encoded this way, the same rule can potentially serve in offline evaluation, regression testing, data filtering, and reinforcement learning.

Cognaptus inference: teams should maintain such constraints as explicit product tests rather than burying them inside broad quality scores. A failure can then be localized to language handling, output protocol, grounding, or another concrete contract.

The image itself can compete with the user

The benchmark treats another failure mode separately: visible text inside an image that contains a competing task.

The setup is deliberately simple. The user asks the model to transcribe text from an image. The image contains clearly readable text that itself looks like an instruction—for example, a request to summarize, calculate, classify, write code, or produce another structured response.

The correct behavior is to transcribe that text, not execute it.

The verifier checks whether the output is sufficiently close to the rendered text using normalized Levenshtein distance, with a threshold below 0.1. This makes the conflict programmatically testable.

Including these instruction-hijacking cases affects models differently. MiMo-VL-7B-RL-2508 moves from 71.40 without the hijacking cases to 64.70 when they are included, while InternVL3.5-8B moves from 70.50 to 64.64. Qwen3.5-35B-A3B is nearly unchanged at 74.60 versus 74.50.

The important result is the heterogeneity. Image understanding is not only about recognizing objects or reading text correctly. Once a model can interpret text inside an image, it also has to decide what authority that text has relative to the user’s instruction.

For document assistants, screenshot agents, interface copilots, and multimodal workflow tools, that deserves its own regression category.

Compliance tests should sit beside capability tests

The paper suggests a useful separation for evaluation pipelines.

Capability tests ask whether the model can solve the underlying task: understand the chart, recognize the object, answer the mathematics question, or extract the document content.

Compliance tests ask whether the model respects the operational contract around that task: correct language, required structure, precise numerical representation, grounding format, and precedence between user instructions and image-embedded text.

The two interact, but the experimental results show why collapsing them into one score is risky.

Cognaptus inference: a multimodal deployment gate should retain both. A new fine-tune can pass a compliance suite and still require independent testing for reasoning, perception, OCR, document understanding, and other underlying capabilities. Conversely, a strong general-purpose model can remain unsuitable for a structured workflow if its output contract is unreliable.

This separation also makes diagnosis cheaper. When a system fails, operators can distinguish “the model did not know the answer” from “the model knew enough but violated the protocol.”

Where the evidence stops

The results are strong within the benchmark’s scope, but that scope is specific.

“Multilingual” here means Chinese and English, not broad many-language coverage. Training experiments use two 8B model families, so the reported learning effects do not establish how substantially different architectures or scales will respond.

The instruction-hijacking test also represents one particular conflict: a genuine OCR request versus a visibly rendered competing task. It establishes that visible image text can interfere with instruction following and provides a reproducible way to measure that behavior. It does not establish resistance to arbitrary multimodal prompt injection, stealth attacks, or every form of adversarial image manipulation.

The broader capability results supply the other important boundary. Qwen’s average rises only from 66.83 to 67.16, while InternVL’s falls from 63.65 to 63.44. The paper reports no repeated-training uncertainty or confidence intervals, so these small changes should not carry more interpretive weight than the experiments support.

There is also a documentation discrepancy worth preserving rather than silently repairing: the paper describes 52 instruction-constraint subcategories, while the rendered appendix table enumerates 40 visible rows.

Measure obedience as obedience

MM-IFEval-Pro makes instruction following more concrete by turning many requirements into executable checks and by testing what happens when text inside an image competes with the user’s request.

Its training results show that this behavior is meaningfully improvable. They also provide a useful warning about what those improvements mean.

A model can become substantially better at satisfying the rules around a task while its broader multimodal capability changes very little.

For operators, that is not a weakness of instruction-following evaluation. It is a reason to measure the property precisely—and to avoid asking one benchmark to stand in for another.

Cognaptus: Automate the Present, Incubate the Future.


  1. Changming Xiao and Zhenliang Ni and Jinhui He and Han Shu and Jie Hu (2026). MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models. arXiv:2609.04859. https://arxiv.org/abs/2609.04859 ↩︎