TL;DR for operators
A clinical LLM can look strong on benchmark accuracy and still behave poorly when the evaluation asks a different question: will a small clinically meaningful change move its answer in the intended direction, and will clinically unchanged evidence remain stable when demographic descriptors are added?
The paper tests those properties separately rather than compressing them into one safety score. Across six models and 150 MedQA questions, overall counterfactual validity was only 25.3%. In a separate demographic stress test, 20.3% of 5,400 baseline-to-prefix comparisons changed the selected answer. Explanation drift and automated stereotype flags added two further screening dimensions.
For model intake or release review, Cognaptus would use these tests as additional gates after conventional accuracy evaluation. Cases with invalid counterfactuals, answer changes, large explanation drift, or serious stereotype signals can be routed to expert review. The evidence supports comparative screening and regression detection. It does not establish clinical deployability, causal demographic bias, or patient harm.
A high score does not tell you how the decision will move
Suppose a clinical AI review has reached a familiar point: several candidate models answer enough medical questions correctly to remain under consideration. Accuracy has done its first job. The harder question is whether similar accuracy implies similar reliability once the input changes slightly.
The results in Beyond Accuracy: Counterfactual Fragility and Demographic Bias in Clinical Evaluation of LLMs1 show why that assumption is unsafe. DeepSeek R1 Distill Qwen 32B achieved 81.6% accuracy but only 10.0% counterfactual validity. MedGemma 27B reached 87.1% accuracy and 63.3% counterfactual validity. Their benchmark scores are both relatively high, but their behavior under minimal case changes is very different.
The six-model profile makes the separation clearer:
| Model | Accuracy | Counterfactual validity | Answer change | Mean EDD | Stereotype flags/question |
|---|---|---|---|---|---|
| GPT-4o mini | 79.4% | 24.7% | 30.7% | 0.244 | 3.55 |
| GLM-4.7 Flash | 71.9% | 10.7% | 48.7% | 0.193 | 5.09 |
| DeepSeek R1 Distill Qwen 32B | 81.6% | 10.0% | 30.0% | 0.181 | 1.81 |
| MiniMax M2.5 | 82.9% | 25.3% | 31.3% | 0.170 | 3.98 |
| MedGemma 27B | 87.1% | 63.3% | 16.0% | 0.169 | 3.68 |
| OpenBioLLM 8B | 55.0% | 18.0% | 59.3% | 0.462 | 2.75 |
The practical message is not that accuracy becomes unimportant. It is that accuracy describes one behavior: whether the model selected the correct answer on the original case. It does not establish that the model has located the decision boundary well enough to react correctly when a clinically relevant fact changes.
Counterfactual validity tests whether the answer moves when it should
The paper adds a local perturbation test. For each model-question pair, the system starts from the model’s original answer, selects a different answer option as the target, and asks the model for the smallest realistic case change that should make that target correct. The altered case is then answered again.
An attempt counts as valid only when the requested target is reached through a plausible, limited modification. Across 900 attempts, only 228 were valid: 25.3%.
The failure taxonomy is useful because it diagnoses why an attempt failed rather than merely reporting the aggregate rate. Of the 672 invalid attempts, 385 were “same direction” failures: the modification did not actually move the clinical case toward the intended alternative. Another 175 changed the case too dramatically, 104 were circular, and eight failed formatting.
That distribution suggests that many failures were not simply parser noise. The dominant problem was failure to cross the intended decision boundary with a sufficiently small change.
This should still be read narrowly. Counterfactual validity is a local consistency test constructed around MedQA cases. It is not evidence that the model possesses a general causal model of medicine.
Robustness also requires staying still when the evidence should stay still
The complementary test asks the opposite question. If the underlying clinical facts are intended to remain unchanged, does the answer remain stable when demographic descriptors are introduced?
The study compares each no-demographic baseline with six race-by-sex prefixes: male or female crossed with White, Black, or Hispanic. Across 5,400 paired comparisons, 1,097 answers changed, a rate of 20.3%.
That number is operationally relevant because it identifies cases where a model’s selected answer is sensitive to the stress-test perturbation. It should not be read as a clean estimate of race or sex effects. The demographic template can conflict with information already present in a vignette, including age, sex, newborn status, and family context.
The correct use is therefore screening: find cases whose answers move under the applied demographic perturbation, then inspect whether that movement has a clinically defensible reason.
Explanations can drift even when answer changes miss the problem
Discrete answer stability does not capture everything. Two variants may produce the same selected option while justifying it differently.
The paper therefore introduces Explanation Demographic Dissonance, or EDD. For each question, it embeds the baseline explanation and six demographic-variant explanations, computes the cosine similarity for all 21 pairs, and defines:
Higher EDD means the explanations are less semantically consistent across variants. OpenBioLLM 8B, for example, had the highest reported mean EDD at 0.462, while MedGemma’s mean was 0.169.
EDD should not be converted into a bias score. It measures explanation drift. A low value does not guarantee the absence of problematic demographic reasoning, while a high value does not establish that the explanation is biased.
That is why the framework keeps a separate automated stereotype screen. Across the evaluation, the judge produced 3,128 stereotype-evidence flags. Most were labeled clinically misleading rather than clinically dangerous. These classifications are useful for triage, but all were generated automatically rather than confirmed by clinicians.
Use the audit as a release filter, not a clinical certificate
The strongest business application is a multidimensional model profile.
What the paper directly shows: accuracy, counterfactual validity, demographic answer stability, explanation drift, and stereotype evidence can vary substantially across models. Treating one metric as a proxy for the others would discard information visible in the reported results.
What Cognaptus infers for operations: model intake should retain conventional accuracy testing but add perturbation-based screening. The same suite can then be rerun after a model update, prompt change, retrieval-system modification, or decoding change. Regression becomes easier to detect because the comparison asks not only whether accuracy moved, but whether local reasoning and demographic robustness changed as well.
The workflow also changes how expert time can be allocated. Instead of asking clinicians to inspect every generated record, an organization can prioritize answer-changing cases, unusually high-EDD records, invalid counterfactuals, and severe stereotype signals.
For that process to remain auditable, the organization should retain the model version, prompt templates, decoding parameters, access dates, raw outputs, parser results, and automated judge traces. Without those records, a later change in safety-screening behavior becomes difficult to reproduce or investigate.
The evidence stops before clinical approval
Several limitations determine how far this framework can be taken.
The study uses 150 deterministically sampled MedQA questions and does not establish specialty-level representativeness. Counterfactual plausibility and stereotype judgments are automated, with no clinician adjudication or inter-rater agreement. The demographic template sometimes inserts a fixed 45-year-old descriptor and can conflict with the original case. No age-sensitivity reruns were conducted to separate these template effects from demographic sensitivity more cleanly.
The six-model comparison is therefore a model-profile exercise, not a population estimate for clinical LLMs. The paper also lacks complete model access and version manifests.
These constraints do not erase the value of the audit. They define it. It can reveal where a model behaves differently under targeted perturbations and where expert review should concentrate. It cannot convert those signals into proof of unfairness, causal understanding, patient harm, or readiness for autonomous clinical use.
Accuracy remains a reasonable first gate. This paper shows why it should not be the last one.
Cognaptus: Automate the Present, Incubate the Future.
-
Chaitai Deb Purkayastha and Bharath Kumar Bolla and Vishnu Surya Reddy Nandi (2026). Beyond Accuracy: Counterfactual Fragility and Demographic Bias in Clinical Evaluation of LLMs. arXiv:2609.32807. https://arxiv.org/abs/2609.32807 ↩︎