The Last Layer Is Not the Last Word on Speech Quality
TL;DR for operators If a speech-quality pipeline extracts features only from the final layer of a foundation model, it may be leaving measurable predictive accuracy on the table. In CAL-MOS, Wav2BERT’s average utterance-level mean squared error falls from 0.466 with the last layer to 0.395 when the strongest intermediate layer is selected; system-level MSE falls from 0.116 to 0.061. ...