Cover image

Noise Rewrote the CT Leaderboard

TL;DR for operators A team using a clean CT leaderboard must decide which high-scoring models deserve costly validation. That ranking may be useful for screening, but it may not survive the operating conditions the model will actually encounter. When the same 200 breast CT cases were exposed to mild, previously unseen Poisson noise, the clean and noisy rankings became essentially uncorrelated, with Spearman’s $\rho$ of about 0.04. The clean-data champion no longer reduced reconstruction error relative to the condition-matched baseline, giving it zero calibrated headroom, while a method ranked lower on clean inputs became the noisy-condition leader. ...

August 6, 2026 · 9 min · Zelina