Cover image

Before You Spend 10 Trillion Tokens: Separate Width From Horizon

TL;DR for operators A training team preparing a multi-trillion-token MoE run usually cannot afford to test several full-scale learning rates. Kim et al. show a way to reduce that search before the expensive run begins: transfer the learning-rate optimum across model width, then estimate separately how that optimum moves as the token budget grows.1 ...

September 23, 2026 · 7 min · Zelina
Cover image

The Right Answer Is Not Enough: Verify the Reasoning Before You Train on It

TL;DR for operators A synthetic reasoning trace can end with the correct answer and still contain intermediate steps you would not want a model to imitate. ORACLE1 addresses that data-quality problem by checking reasoning one step at a time: it uses a symbolic reasoning engine when a step can be formalized, and LLM-based correctness and feasibility judgments when it cannot. ...

September 17, 2026 · 7 min · Zelina
Cover image

Alignment Is a Coverage Problem Before It Is a Loss-Function Problem

TL;DR for operators An alignment team with a fixed preference dataset faces a deceptively simple decision: train offline with a direct method, or spend more compute to keep generating and evaluating new responses during training. The cheaper route is not always the safer one. The survey by Tarun Raheja and Nilay Pochhi1 highlights a theoretical coverage result under which offline contrastive preference learning needs stronger coverage of possible responses than online reinforcement learning. If useful responses lie outside the regions represented in the fixed dataset, an offline learner has no direct learning signal there. Online methods can generate new data and therefore operate under a weaker, partial-coverage requirement. ...

September 15, 2026 · 8 min · Zelina
Cover image

Compress the Activation, Spend the Memory on Rank

TL;DR for operators A LoRA fine-tuning job can fit its trainable parameters comfortably on a GPU and still run out of memory because backpropagation retains large intermediate activations. CARE-LoRA attacks that remaining buffer rather than shrinking the adapter itself. Zhang et al. store LoRA’s already-compressed activation plus a small reconstruction matrix, then use those tensors to approximate only the gradient needed for one of LoRA’s two trainable matrices.1 ...

September 5, 2026 · 7 min · Zelina
Cover image

Route Cause Analysis: Stop Sending Every AI Failure to Training

TL;DR for operators An AI failure is an observation, not a diagnosis. A low benchmark score, an incorrect answer, or a broken agent run does not tell you whether the underlying problem belongs in the training data, model objective, retrieval policy, procedural instructions, tool interface, or execution environment. Treating all of these as “model quality” produces expensive interventions with weak causal logic. ...

July 22, 2026 · 19 min · Zelina
Cover image

Stale Gradients, Fresh Economics: CoCD’s Lightweight Route to Zeroth-Order AI

Memory is usually treated as a luxury in machine learning. More parameters, more activations, more optimiser state, more logs, more everything. Then the invoice arrives, the device overheats, and someone rediscovers the ancient corporate virtue of not wasting things. The paper Turning Stale Gradients into Stable Gradients makes a modest but interesting proposal: perhaps an optimiser should not throw away old gradient information just because it is old.1 In the right setting, yesterday’s partial derivative is not spoiled milk. It is a slightly outdated map. If the terrain has not shifted too violently, it may still point in a useful direction. ...

June 13, 2026 · 16 min · Zelina
Cover image

High Entropy, Low Drama: The Internal Fingerprint of LLM Reasoning

Debugging a reasoning model usually starts at the wrong end. A model gives a wrong mathematical answer, so we inspect the final output. Then we inspect the chain-of-thought. Then we compare benchmark scores, sample more answers, compute pass rates, and hope the model’s visible reasoning trace tells us what happened inside. This is convenient. It is also a little like diagnosing a factory by reading only the shipping label. ...

May 31, 2026 · 15 min · Zelina
Cover image

RL Needs a Menu, Not a Miracle

RL Needs a Menu, Not a Miracle Menus are underrated. When a language model knows only one way to solve a problem, reinforcement learning can mostly reward or punish that route. It can make the model more confident, more selective, and sometimes more verbose. But it has little room to choose among genuinely different ways of reaching the answer. ...

May 25, 2026 · 14 min · Zelina
Cover image

When Reasoning Pays (and When It Cheats): Fixing RL Signals in LLM Training

Scorecards are useful until people learn how the scorecard works. That is not a cynical observation. It is basic management. Sales teams optimize for commission rules. Customer-service teams optimize for handle-time dashboards. Students optimize for exams. And language models, with their charming lack of shame, optimize whatever reward function we put in front of them. ...

March 30, 2026 · 17 min · Zelina
Cover image

Training Models to Explain Themselves: Counterfactuals as a First-Class Objective

Rejected. That is where counterfactual explanations usually enter the story. A loan applicant is declined by an automated system. A hiring candidate is filtered out. An insurance customer is priced into an unfavorable category. The counterfactual explanation is supposed to answer a practical question: what would need to change for the model to give me the desired outcome? ...

January 24, 2026 · 16 min · Zelina