Cover image

The Recipe Moves With the Run

TL;DR for operators A small-scale hyperparameter sweep can narrow the search for a larger training run, but the resulting recipe is conditional on more than model size. In the OpenEuroLLM experiments, the loss-optimal batch size increased with model size and token budget, while the learning-rate relationship changed depending on whether batch size was optimized jointly or fixed by infrastructure. ...

September 12, 2026 · 8 min · Zelina
Cover image

Same Algorithm, Different Outcome: What 33,000 Actor-Critic Runs Reveal

TL;DR for operators Two reinforcement-learning systems can use the same named algorithm, train on the same task, and still produce materially different outcomes across random seeds. Shah, Zhu, White, and White investigate why by decomposing actor-critic systems into their lower-level choices rather than treating PPO, SAC, DDPG, or MPO as indivisible packages.1 ...

August 16, 2026 · 8 min · Zelina