Cover image

The Leaderboard Is Not a Clinical Clearance

TL;DR for operators A healthcare team choosing a model, configuration, and safeguards for diagnostic support should treat leaderboard leadership as evidence of capability—not proof of clinical readiness. GPT-5-medium led ClinMM-Bench, yet produced a completely correct diagnosis in only 33.88% of cases. The benchmark tests a difficult but essential requirement: combining clinical details and images as evidence unfolds, revising earlier conclusions, and explaining a diagnosis without omitting decisive information or introducing unsupported claims. Medical specialization and explicit reasoning settings do not improve these abilities consistently across model scales, metrics, or specialties. ...

August 6, 2026 · 8 min · Zelina
Cover image

When X-Rays Talk Back: Grounding AI Diagnosis in Evidence, Not Eloquence

Chest X-rays are not mysterious objects. They are images that radiologists interrogate through a disciplined sequence: find the anatomy, measure what matters, compare against criteria, and then make a diagnostic judgment. The modern vision-language model often skips the middle of that sequence. It looks at the image, produces a polished explanation, and hopes the reader will not ask too aggressively where the evidence came from. This is how medical AI becomes impressive in a demo and uncomfortable in a clinic. Fluency is cheap. Verifiability is expensive. ...

February 27, 2026 · 14 min · Zelina
Cover image

Charting a Better Bedside: When Agentic RL Teaches RAG to Diagnose

TL;DR for operators Diagnosis is not a search-box problem. A clinician does not simply type a symptom list, read a guideline, and pick a disease like ordering takeaway. The useful work is iterative: form a hypothesis, compare against similar cases, notice what does not fit, retrieve again, ignore plausible-looking rubbish, and only then commit. ...

August 24, 2025 · 18 min · Zelina