Cover image

Sparse Routing May Buy You Inspectability, Not Just Efficiency

TL;DR for operators Herbst, Wermter, and Lee find that the analyzed Mixture-of-Experts models often represent tested concepts in far fewer neurons than comparable dense transformers.1 The difference is largest under the hardest probe constraint: when only one neuron is available, MoE experts often approach their own best probe performance while dense feed-forward layers need more dimensions. Models with sparser routing also tend to show cleaner representations. ...

September 11, 2026 · 8 min · Zelina