Cover image

A Correct Answer Can Still Be a Fragile One

TL;DR for operators A clinical LLM can look strong on benchmark accuracy and still behave poorly when the evaluation asks a different question: will a small clinically meaningful change move its answer in the intended direction, and will clinically unchanged evidence remain stable when demographic descriptors are added? The paper tests those properties separately rather than compressing them into one safety score. Across six models and 150 MedQA questions, overall counterfactual validity was only 25.3%. In a separate demographic stress test, 20.3% of 5,400 baseline-to-prefix comparisons changed the selected answer. Explanation drift and automated stereotype flags added two further screening dimensions. ...

October 11, 2026 · 7 min · Zelina