Cover image

When Worse Inputs Score Better: Audit the Credibility Behind the Benchmark

TL;DR for operators A benchmark score can be high without being equally trustworthy as a measure of generalization. In this study, the researchers deliberately degraded benchmark questions before they reached the answering model. At a noisy-router count of eight, 10 of the 12 evaluated models nevertheless scored above their own clean baseline. At nine routers, eight models still did so, and the mean positive excess among above-baseline cases reached 0.086. ...

September 10, 2026 · 7 min · Zelina