Sparse Routing May Buy You Inspectability, Not Just Efficiency
TL;DR for operators Herbst, Wermter, and Lee find that the analyzed Mixture-of-Experts models often represent tested concepts in far fewer neurons than comparable dense transformers.1 The difference is largest under the hardest probe constraint: when only one neuron is available, MoE experts often approach their own best probe performance while dense feed-forward layers need more dimensions. Models with sparser routing also tend to show cleaner representations. ...