Cover image

The Scaling Law Got a Data Manager

TL;DR for operators A useful scaling law does not merely say “bigger is better.” That is not a law; that is a purchasing department with a GPU account. The paper behind this article studies whether the composition of pretraining data can change the compute-optimal balance between model size and downstream data in jet classification.1 The answer, in this setting, is yes. Training from scratch on JetClass produces a nearly balanced scaling rule: as compute grows, the optimal model size and dataset size grow at roughly similar rates. But after pretraining on a JetClass-II corpus augmented with Beyond Standard Model resonance decays, the compute-optimal rule shifts sharply toward downstream data. More of the next compute budget should be spent processing more examples, not inflating the model. ...

June 22, 2026 · 16 min · Zelina
Cover image

The AI Buffet: Why One Supermodel Might Rule the Menu, But Specialty Dishes Still Sell

TL;DR for operators The AI market is not choosing between “one model to rule them all” and “a thousand specialist flowers blooming politely in a procurement spreadsheet.” It is choosing by workload. GPT-4o’s native image generation matters because it folds visual production into the same conversational workspace where users already brainstorm, rewrite, code, and revise. That is not just a model upgrade. It is a distribution upgrade. The GPT-4o system card describes an omni model trained across text, vision, and audio, with stronger multimodal capability and lower API cost than GPT-4 Turbo in OpenAI’s own framing.1 OpenAI’s March 2025 image-generation release then pushed that logic into visual work: generate, critique, revise, and regenerate without leaving the chat.2 ...

April 8, 2025 · 12 min · Zelina