When AUROC Agrees and the Decision Still Changes
TL;DR for operators CRS-Bench1 tests a decision that clean-test AUROC does not fully answer: which pretrained medical image encoder deserves the next round of engineering, labeling, adaptation, and validation investment. Across 15 encoder families, AUROC and the benchmark’s broader reliability score are positively associated, yet 21 of 105 pairwise model choices reverse. The disagreement is not evidence that AUROC is useless. It shows that discrimination can preserve the broad ordering while missing differences in calibration, label efficiency, and robustness that change actual selection decisions. ...