When More Pseudo-Labels Stop Helping: The Active Bottleneck in Weak-to-Strong Learning
TL;DR for operators After building a teacher model, a development team faces a familiar allocation question: generate more teacher-labeled examples, increase student capacity, or do both. The usual scaling intuition is incomplete. Zeinalpour and Najafi show that a student can outperform its teacher even when both use the same linear hypothesis class and training rule, converge exactly, and use neither ridge regularization nor early stopping.1 The improvement comes from a finite Stage-II resource: the student cannot reproduce every component of the teacher, and the resulting restriction can remove more teacher error than useful signal. ...