Cover image

More FLOPs, Worse Choice: Sparse MoE Scaling Has to Price the Cluster

TL;DR for operators Suppose a team has already bought the cluster time: 32 B200 nodes for 20 days. It might seem sensible to prefer the sparse-MoE design that manages to execute the most model FLOPs during that window. In the paper’s reported search, that rule would pick the wrong configuration. The loss-optimal candidate uses about $1.23\times10^{23}$ model FLOPs—roughly 36% fewer than the candidate that realizes about $1.94\times10^{23}$—yet reaches the lower predicted loss. ...

September 11, 2026 · 7 min · Zelina