Before You Spend 10 Trillion Tokens: Separate Width From Horizon
TL;DR for operators A training team preparing a multi-trillion-token MoE run usually cannot afford to test several full-scale learning rates. Kim et al. show a way to reduce that search before the expensive run begins: transfer the learning-rate optimum across model width, then estimate separately how that optimum moves as the token budget grows.1 ...