When the Next Sparse Billion Should Go to Memory, Not Experts
TL;DR for operators A sparse-model team with a fixed marginal parameter budget should not assume that the next increment belongs in more experts. In the LongCat-Flash experiments reported by Liu et al.,1 parameter-equivalent expert scaling performs better earlier, but N-gram embedding scaling takes the lead after the MoE reaches a sufficiently high-sparsity regime. ...