TL;DR for operators
A healthcare team choosing a model, configuration, and safeguards for diagnostic support should treat leaderboard leadership as evidence of capability—not proof of clinical readiness. GPT-5-medium led ClinMM-Bench, yet produced a completely correct diagnosis in only 33.88% of cases.
The benchmark tests a difficult but essential requirement: combining clinical details and images as evidence unfolds, revising earlier conclusions, and explaining a diagnosis without omitting decisive information or introducing unsupported claims. Medical specialization and explicit reasoning settings do not improve these abilities consistently across model scales, metrics, or specialties.
Use the benchmark to compare exact configurations within the intended specialty, identify failure modes, and define review and escalation controls. Procurement, pilot scope, and safety design should follow measured behavior—not aggregate rank, product labels, or the appearance of detailed reasoning.
A winning model still leaves most difficult cases unresolved
A hospital innovation team can review several impressive medical-AI scores and still face an incomplete decision. The team must choose a model, decide which specialty should enter a pilot, determine what evidence the system may handle, and specify when a clinician must take over. A single aggregate accuracy number does not answer those questions.
The central friction is visible in the leading result. GPT-5-medium ranked first, yet its diagnosis was completely correct in only 33.88% of cases. Most outputs were partially rather than fully correct. A model can appear directionally competent while still missing the integration or revision required for a defensible final diagnosis.
ClinMM-Bench was designed to expose this gap.1 It contains 1,089 challenging clinical cases and 3,760 medical images across eight specialties. Information is disclosed over an average of 5.45 dialogue rounds, rather than presented as one complete prompt. The model must carry earlier findings forward as later text and images change the diagnostic picture.
This setting better reflects diagnosis than static medical question answering, but it remains controlled. Evidence is disclosed passively; the model does not choose questions, examinations, or tests. The benchmark tests use of evolving evidence, not management of an actual workup.
The benchmark tests the path to an answer, not only the answer
Final-answer scoring remains part of the evaluation. Two independent judge models—GPT-5-medium and Claude-4.5-Sonnet—score diagnoses from 0 to 2, and the paper averages their judgments. GPT-5-medium achieved the highest mean score at 1.140, followed by Gemini 3 Pro at 1.038 and GPT-5-minimal at 1.031. The best open-weight model, Qwen3-VL-32B, reached 0.718 and produced completely correct diagnoses in 11.20% of cases.
These main comparisons show a substantial gap between evaluated proprietary and open-weight models under a common protocol. They do not show that ownership causes better performance, because it is bundled with differences in data, architecture, scale, training, and inference.
ClinMM-Bench then evaluates the reasoning itself by breaking both model and reference explanations into atomic clinical facts. This step addresses a practical problem: a plausible final diagnosis can conceal an incomplete or unreliable evidentiary path.
The benchmark reports three complementary measures:
| Measure | What it captures | Operational risk if weak |
|---|---|---|
| Fact recall | How much relevant reference evidence the reasoning includes | Important findings may be omitted from the justification |
| Hallucination | How much unsupported clinical content the model introduces | Clinicians may be asked to evaluate invented findings |
| Fact density | How much supported information appears relative to the length of the reasoning | Long explanations may consume review time without adding evidence |
No model dominated all three. GPT-5-medium had the highest fact recall. Gemini 3 Pro had the lowest hallucination score. LLaMA-4-Scout had the highest fact density. Diagnostic accuracy and reasoning quality therefore represent different selection criteria, not interchangeable proxies.
For a clinical-support pilot, this changes the evaluation target. A model that retrieves more relevant facts may still add unsupported ones. A concise model may be efficient to review but omit decisive evidence. A high aggregate diagnostic score may still hide failure in the specialty or image type that the organization intends to deploy.
Specialization and reasoning mode require configuration-specific evidence
Two product features often appear reassuring in model selection: medical specialization and explicit reasoning. The paper’s paired comparisons show why neither should be accepted as a default advantage.
Medical specialization improved some smaller-model diagnostic results and reduced hallucination in Gemma–MedGemma comparisons. The accuracy benefit, however, was not consistent at larger scale. This is not evidence that domain adaptation is ineffective. It shows that its benefit depends on model size, metric, and task condition. A procurement team should test the exact specialized model it plans to use rather than infer performance from the label “medical.”
The reasoning-mode comparison is more direct. Across tested Qwen3-VL models at 4B, 8B, and 32B scales, reasoning-mode variants did not consistently improve diagnostic accuracy. They also showed lower fact recall and higher hallucination than corresponding non-reasoning variants. These are controlled descriptive comparisons within one model family, not a universal result about all reasoning systems.
The likely product mistake is to equate a longer or explicitly enabled reasoning trace with better clinical reasoning. ClinMM-Bench measures whether the trace preserves relevant facts and avoids unsupported ones. Under that test, more visible reasoning can create more surface area for omission, drift, or fabrication.
The appropriate deployment response is regression testing. Every change in model version, prompt, reasoning setting, image preprocessing, or inference configuration should be rerun against the target specialty. Feature names are not substitutes for measured behavior.
Five failure modes imply five different controls
The paper’s qualitative analysis classifies representative errors into information synthesis failure, knowledge mapping error, perception error, premature closure, and visual hallucination. This taxonomy is exploratory rather than a quantified prevalence estimate, but it is operationally useful because each category points to a different safeguard.
| Failure mode | What fails | Candidate control for a pilot |
|---|---|---|
| Information synthesis | Relevant evidence is retained but not combined across turns | Structured evidence ledger and cross-turn contradiction checks |
| Knowledge mapping | Observations are connected to the wrong clinical mechanism or diagnosis | Retrieval support, differential-diagnosis constraints, specialist review |
| Perception error | The model misreads the image or misses a visible feature | Image-quality gates and independent visual verification |
| Premature closure | An early hypothesis persists despite conflicting later evidence | Mandatory hypothesis revision after material new findings |
| Visual hallucination | The reasoning cites image findings that are not present | Grounding checks requiring localization or explicit image evidence |
These controls are Cognaptus inferences from the failure taxonomy, not interventions tested by the paper. Their value is in turning an undifferentiated “model error” into an auditable workflow condition.
The affected user is not only the clinician. Procurement teams need specialty comparisons, safety teams need escalation triggers, and product teams need regression cases. Clinical leaders must define whether the model may suggest a differential, rank alternatives, draft a rationale, or present a final diagnosis. One result may support a narrow use while disqualifying a broader one.
Specialty averages should determine pilot scope
Aggregate results conceal substantial specialty variation. Across models, diagnostic accuracy was highest in Neurology and Ophthalmology and lowest in Radiology and Internal Medicine. That variation alone argues against an institution-wide model choice based on one overall score.
The benchmark composition also requires care. Radiology contributes 679 of the 1,089 cases, or 62.35% of the dataset. Internal Medicine and Emergency Medicine each contribute 23 cases. The paper reports bootstrap confidence intervals for specialty-level results, but smaller specialty samples still provide less stable operational evidence than the overall benchmark.
A model-selection team should define the intended specialty, modality, case complexity, acceptable omission and hallucination rates, and human-review capacity before reading the leaderboard. It can then judge whether performance and review burden fit the workflow.
This is a better use of ClinMM-Bench than declaring one model “best for medicine.” The benchmark supports comparative selection under its protocol. It does not supply a universal clinical ranking.
What the benchmark cannot establish
The cases are drawn from challenging published reports, including uncommon and difficult presentations. Performance should not be interpreted as an estimate of routine diagnostic accuracy or disease prevalence. The benchmark is deliberately a stress test, not a representative clinical sample.
Its evaluation pipeline also relies extensively on language models for case validation, conversion, quality control, diagnosis judging, and atomic-fact analysis. Multi-stage checks and sampled expert validation reduce measurement risk, but they cannot eliminate shared evaluator bias. Expert review does not cover every retained case.
The study does not test prospective deployment, patient outcomes, workflow integration, time pressure, or clinicians’ responses to model advice. It also does not establish causal effects of scale, specialization, reasoning mode, or ownership.
These boundaries specify its proper role: screening configurations before higher-cost clinical validation, diagnosing recurring failure patterns, and detecting regressions after system changes.
Clinical readiness begins after the leaderboard
ClinMM-Bench converts difficult case reports into progressive multimodal dialogues and evaluates both the diagnosis and the evidentiary quality of its reasoning. It provides a clearer account of what current systems can and cannot be trusted to do.
For operators, the selection rule should be explicit: choose by intended specialty, verify the exact model and inference configuration, measure evidence coverage and hallucination separately from answer accuracy, and connect known failure modes to review and escalation controls.
A first-place score can justify further evaluation. It cannot serve as clinical clearance.
Cognaptus: Automate the Present, Incubate the Future.
-
Rui Yang and Weihao Xuan and Yi Lin and Zhuhan Bao and Jonathan Chong Kai Liew and Matthew Yu Heng Wong and Nicolás Lescano and Nikita R. Paripati and Emily Ling-Lin Pai and Jiarui Liu and Heli Qi and Heng-Jui Chang and Benny Kai Guo Loo and Huitao Li and Kunyu Yu and Yufan Wang and Chuan Hong and Shijian Lu and Douglas Teodoro and Naoto Yokoya and Ross Koppel and Mona Diab and Hua Xu and David W. Bates and Nan Liu and Yifan Peng (2026). Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases. arXiv:2607.25933. https://arxiv.org/abs/2607.25933 ↩︎