Cover image

Parameter Matching Is Not Baseline Matching

TL;DR for operators A team choosing between two tiny models for an edge product could reasonably treat the same nominal parameter budget as a fair baseline. In this controlled study at roughly 60K parameters, however, five approximately parameter-matched full-precision transformer shapes span validation losses from 2.835 to 3.477 nats per byte at the 16M-byte training budget—a 22.6% difference produced by the tested shape package alone. Against the weakest transformer shape, the model that dynamically distributes computation across several internal pathways looks roughly 19–21% better; against the strongest full-precision transformer, it is only 0.95% better, inside the study’s seed-noise threshold. ...

October 10, 2026 · 7 min · Zelina