Before the Model Speaks, Check Whether the Audio Does
TL;DR for operators A speech model can perform well when given valid speech and still behave badly on audio that should never have reached generation. Mengzhe Geng’s SURE-Voice study1 isolates that earlier decision. On the 480-example held-out SURE-Extended test, six raw speech/audio LLMs achieved only 0.000 to 0.133 accuracy on unsupported inputs. Placing the same fixed audio gate in front of each model raised unsupported accuracy to 0.919, preserved each backbone’s stored supported accuracy of 0.919 to 0.970, and reduced downstream model calls from 480 to 287. ...