Cover image

Same Modalities, Different Loss: Pairing Is Its Own Multimodal Data Budget

TL;DR for operators A multimodal dataset can contain the same amount of image, text, audio, or other modality-specific data yet produce different model performance depending on how often those modalities occur together. Marcus Ma and Shrikanth Narayanan isolate that distinction in Paired Multimodal Scaling Laws1, varying pairing while holding modality-specific data budgets fixed across three supervised environments. ...

October 5, 2026 · 7 min · Zelina
Cover image

More Data, Better Learner: What Children Reveal About Learning Efficiency

TL;DR for operators Adding more data and becoming better at learning from data are different objectives. Across five longitudinal vocabulary datasets covering American English, Norwegian, and Japanese, young children show strongly increasing returns to developmental experience. Using a matched per-word estimator, the language models examined in the study produce a median acceleration estimate of 1.16, with an interquartile range of 0.93–1.46. Corrected child estimates fall around 10.4–13.8. ...

September 24, 2026 · 7 min · Zelina
Cover image

Before You Spend 10 Trillion Tokens: Separate Width From Horizon

TL;DR for operators A training team preparing a multi-trillion-token MoE run usually cannot afford to test several full-scale learning rates. Kim et al. show a way to reduce that search before the expensive run begins: transfer the learning-rate optimum across model width, then estimate separately how that optimum moves as the token budget grows.1 ...

September 23, 2026 · 7 min · Zelina
Cover image

Scale the Split: MoE Compute Allocation Should Move With the Budget

TL;DR for operators When a sparse Mixture-of-Experts model receives a larger training budget, the additional computation has at least two architectural destinations: attention capacity and expert feed-forward capacity. Treating their existing balance as fixed can leave the larger model internally misallocated. Li et al. show this experimentally in GPT-style sparse MoE models.1 At fixed compute and sparsity, varying the expert-versus-attention allocation produces a clear loss minimum. As total compute increases, the loss-minimizing allocation shifts toward more expert computation, but the rate of that shift depends on sparsity. ...

September 12, 2026 · 7 min · Zelina
Cover image

The Recipe Moves With the Run

TL;DR for operators A small-scale hyperparameter sweep can narrow the search for a larger training run, but the resulting recipe is conditional on more than model size. In the OpenEuroLLM experiments, the loss-optimal batch size increased with model size and token budget, while the learning-rate relationship changed depending on whether batch size was optimized jointly or fixed by infrastructure. ...

September 12, 2026 · 8 min · Zelina
Cover image

More FLOPs, Worse Choice: Sparse MoE Scaling Has to Price the Cluster

TL;DR for operators Suppose a team has already bought the cluster time: 32 B200 nodes for 20 days. It might seem sensible to prefer the sparse-MoE design that manages to execute the most model FLOPs during that window. In the paper’s reported search, that rule would pick the wrong configuration. The loss-optimal candidate uses about $1.23\times10^{23}$ model FLOPs—roughly 36% fewer than the candidate that realizes about $1.94\times10^{23}$—yet reaches the lower predicted loss. ...

September 11, 2026 · 7 min · Zelina
Cover image

The Scaling Law Got a Data Manager

TL;DR for operators A useful scaling law does not merely say “bigger is better.” That is not a law; that is a purchasing department with a GPU account. The paper behind this article studies whether the composition of pretraining data can change the compute-optimal balance between model size and downstream data in jet classification.1 The answer, in this setting, is yes. Training from scratch on JetClass produces a nearly balanced scaling rule: as compute grows, the optimal model size and dataset size grow at roughly similar rates. But after pretraining on a JetClass-II corpus augmented with Beyond Standard Model resonance decays, the compute-optimal rule shifts sharply toward downstream data. More of the next compute budget should be spent processing more examples, not inflating the model. ...

June 22, 2026 · 16 min · Zelina
Cover image

The Viscosity Budget: Why Softmax Is Not Just a Knob

TL;DR for operators A new paper by Jose Marie Antonio Miñoza, Erika Fille T. Legara, and Christopher P. Monterola argues that a log-sum-exp neural layer is not merely analogous to a viscous Hamilton-Jacobi equation. Under the paper’s parameterisation, it is exactly the Hopf-Cole solution of one, evaluated at the input point.1 The operational point is not “neural networks are physics now”, although someone will certainly try to put that on a slide. The point is cleaner: one parameter, $\varepsilon$, simultaneously controls softmax temperature, PDE viscosity, and entropy-regularised convex optimisation. That makes smoothness, expressiveness, robustness, attribution sharpness, and scaling behaviour mathematically coupled. ...

June 18, 2026 · 18 min · Zelina
Cover image

One Pass to Forecast Them All: Toto 2.0 and the Scaling Recipe for Time-Series AI

Forecasting is where machine learning often learns humility. A language model can sound clever while being wrong. A forecasting model has fewer hiding places. Revenue arrives or it does not. CPU saturation happens or it does not. Demand spikes, latency drifts, inventories rot, turbines fail, and the spreadsheet smiles politely before punishing everyone involved. This is why time-series foundation models have been treated with a particular kind of suspicion: useful, interesting, sometimes impressive, but not yet comfortably scalable in the way large language models became scalable. ...

June 5, 2026 · 18 min · Zelina
Cover image

Filter Bubble Bursts: When Common Crawl Beats Clean Data

Cleaning is comforting. Every serious AI team has some version of the same ritual. Remove spam. Remove repetition. Remove bad language detection. Remove low-quality pages. Remove documents that look too weird, too short, too duplicated, too uneducational, too internet. Then hope the model learns from the respectable leftovers. That instinct is not foolish. In small or compute-constrained training runs, filtering often helps. The expensive mistake is treating that local truth as a permanent law. ...

June 4, 2026 · 14 min · Zelina