Cover image

Following Instructions Is Not the Same as Knowing More

TL;DR for operators A multimodal model can become much better at obeying instructions without becoming much better at the underlying tasks those instructions govern. In the experiments examined here, one 8B vision-language model gains 10.58 percentage points on a targeted instruction-following benchmark and another gains 22.91 points. Yet their average results across broader STEM, VQA, OCR, and document-understanding tests move by only +0.33 and -0.21 points. ...

October 1, 2026 · 7 min · Zelina
Cover image

The Grader Is Part of the System: What RLVR Verifier Audits Reveal

TL;DR for operators A mathematical verifier is not an invariant scoring utility. Esther Xin’s Where the Verifier Fails1 audits four implementations and configurations using 307,420 verifier verdicts and finds acceptance of mathematically equivalent answers ranging from 53.8% to 95.2%. Two extraction configurations of the same math-verify library disagree on 49.9% of equivalent cases where both return a verdict. ...

September 29, 2026 · 7 min · Zelina
Cover image

The KL You Weren’t Watching: Moving RLVR Stability to the Query Side

TL;DR for operators RLVR teams face a familiar trade-off: stronger constraints can stabilize training, but constraining the response distribution too tightly can also suppress useful exploration. The paper identifies a second stability problem that response-side KL alone can miss. The same model being optimized also assigns probabilities to the training queries themselves, and those probabilities can shift substantially even when the dataset stays fixed. ...

September 28, 2026 · 8 min · Zelina
Cover image

Check Your Work: Why Self-Verification Deserves Its Own Training Budget

TL;DR for operators A post-training team deciding where to spend its next training budget should not infer verification ability from task accuracy. In Learning to Self-Verify Makes Language Models Better Reasoners, Chen et al. find that training models to solve mathematical problems better does not reliably make them better at judging whether solutions are correct.1 Training the reverse capability behaves differently: models trained only to judge their own generated solutions subsequently solve problems about as well as models trained directly for generation. ...

September 3, 2026 · 8 min · Zelina
Cover image

Synthetic Experience, Real Transfer: Build the Test Before You Scale the Data

TL;DR for operators Synthetic data should not be budgeted as a cheaper substitute for human examples. It should be treated as infrastructure for producing controlled training experience. The operational sequence is: generate tasks that can actually be executed and scored; verify and repair them before spending compute on trajectories; choose a training objective that reinforces the capability you want rather than merely reproducing successful-looking behavior; and test the resulting model outside the environment in which that experience was generated. ...

September 3, 2026 · 8 min · Zelina
Cover image

When the Simulator Becomes the Curriculum

TL;DR for operators A simulator can generate effectively unlimited trajectories. That does not automatically make those trajectories useful training data for a reasoning model. Sim2Reason turns simulated mechanics into questions with answers that can be checked automatically, then uses those questions for reinforcement-learning post-training. The reported transfer is substantial enough to matter operationally: on Qwen2.5-32B, International Physics Olympiad mechanics accuracy rose from 19.8% to 25.2%. By contrast, supervised fine-tuning on 200,000 teacher-generated trajectories lowered the same score to 15.9%. ...

September 2, 2026 · 7 min · Zelina
Cover image

Compile Once, Train Later: Offline RL Moves Code-Model Verification Upstream

Compile Once, Train Later: Offline RL Moves Code-Model Verification Upstream Code assistants have a small accounting problem. Not the glamorous kind involving model capability, agentic workflows, or yet another dashboard with a glowing neural blob. The ordinary kind: every time a model proposes code during reinforcement learning, someone—or something—has to run it, test it, score it, and feed that score back into training. ...

June 3, 2026 · 14 min · Zelina
Cover image

Judge Math-Not by Its Parser

Opening — Why this matters now The AI industry has discovered a wonderfully pedestrian way to misread progress: build models that can solve harder math problems, then grade them with evaluators that panic when 2040 minutes is not written as 34 hours. That is not a joke. It is the central irritation behind “Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity”, an arXiv paper that examines how mathematical reasoning benchmarks can be distorted by rigid symbolic verification.1 ...

April 27, 2026 · 12 min · Zelina
Cover image

When RL Needs a Tour Guide: OGER and the Business of Smarter Exploration

Training a reasoning model is starting to look less like feeding a student more textbooks and more like taking that student into a difficult city with a very opinionated guide. The guide should not carry the student through every street. That creates a tourist, not a navigator. But leaving the student alone with a reward signal that says only “correct” or “wrong” is not exactly enlightened pedagogy either. The student may find one narrow route, repeat it forever, and call that intelligence. We have all seen corporate training programs with roughly this level of imagination. ...

April 23, 2026 · 18 min · Zelina
Cover image

The Data Diet for Reasoning Models: Why Less (But Smarter) Wins

A model-training team has a familiar bad habit: when the model fails, it asks for more. More examples. More domains. More synthetic prompts. More compute. More benchmarks to average over until the unpleasant details become small enough to ignore. This habit is understandable. It is also expensive. And, according to SuperNova, it may be the wrong first instinct. ...

April 10, 2026 · 16 min · Zelina