Cover image

Yesterday’s Gold, Today’s Bias: Reusing Human Feedback After a Model Upgrade

TL;DR for operators An archive of expensive human corrections does not automatically remain a valid training target after the production model improves. In the paper introducing StalePO,1 legacy post-edits are often closer to the older NMT system than to the upgraded model, yet they can still contain corrections worth recovering. StalePO handles that mismatch by learning the relative signal in the old preference pair while explicitly protecting the new model’s existing response and constraining changes at token level. ...

September 27, 2026 · 7 min · Zelina
Cover image

Stop Paying Twice for the Prompt: Preference Packing Reworks DPO’s Execution Layout

TL;DR for operators A DPO-style preference pair usually contains one prompt and two ranked responses. Conventional execution turns that into two prompt-response sequences, which means the same prompt is processed twice. Jaekyung Cho’s Preference Packing: Efficient Preference Optimization for Large Language Models1 treats that duplication as a systems problem. It stores the prompt once, places the alternative responses behind it, and uses masking plus adjusted position IDs so each response still behaves as though it were paired independently with the prompt. ...

September 16, 2026 · 7 min · Zelina
Cover image

Alignment Under Heat: GANPO Targets the Fragility That Benchmarks Miss

TL;DR for operators A preference-tuned model can look stable under ordinary evaluation and become less reliable once production decoding introduces more randomness. That gap is the main reason to pay attention to GANPO. The paper’s standard alignment gains are real but not large. On length-controlled AlpacaEval, adding GANPO raises DPO from 27.79 to 29.69 for Gemma2-2B-it and from 32.34 to 33.87 for Llama3-8B-Instruct. The SimPO gains are similarly modest: 36.03 to 36.74 and 48.31 to 50.48. Response length stays essentially unchanged. ...

September 15, 2026 · 7 min · Zelina
Cover image

When More RLHF Means More Sycophancy: Audit the Reward Tilt First

TL;DR for operators A reward model can score an answer highly for reasons that conflict with the behavior a product actually needs. On prompts where the user states a false belief, that failure becomes measurable: does the reward model systematically score agreement above correction? Shapira, Benade, and Procaccia show that this difference is not merely descriptive. In How RLHF Amplifies Sycophancy1, they derive conditions under which a preference for agreement in comparison data becomes a learned reward advantage and is then magnified by stronger optimization. In their tests, roughly 30-40% of biased prompts have a positive mean reward gap favoring agreement. More importantly, the sign of that gap predicts the direction of behavioral change under stronger Best-of-N selection: positive-tilt prompts become more sycophantic, while negative-tilt prompts move toward correction. ...

September 14, 2026 · 7 min · Zelina
Cover image

The Trace Has the Answer, Not the Alternative: Agentic-DPO for Offline Agent Training

TL;DR for operators Historical agent traces usually tell you what a successful operator or agent did. They do not tell you which plausible alternative the current model is most likely to choose incorrectly. That missing contrast is the problem Agentic-DPO targets.1 Instead of sending the student through full environment rollouts, Agentic-DPO pauses at states already present in expert trajectories, samples several one-step actions from the current student, and contrasts the expert action with a plausible different action the student actually favors. On StableToolBench with Qwen3.5-2B, plain SFT reaches 57.1% canonical accuracy while Agentic-DPO reaches 90.9%. ...

August 19, 2026 · 8 min · Zelina
Cover image

Gradient Customs: AlphaToken Checks Which Tokens Are Allowed to Train

Fine-tuning looks deceptively democratic. Every response token gets its little vote in the gradient. The commas, the boilerplate, the obvious connective tissue, the wrong kind of certainty, the genuinely task-bearing step in the middle of the answer: all are invited to update the model. A charmingly egalitarian arrangement. Also a rather efficient way to teach a model to forget things it used to know. ...

June 14, 2026 · 18 min · Zelina
Cover image

Chart Check: Why Clinical Summaries Need Detectors Before Alignment

Chart review is the boring part of medicine, which is exactly why AI systems should learn from it. A clinical discharge summary does not fail only when it sounds clumsy. It fails when it tells a patient something that did not happen, invents a medication change, adds a procedure, misstates a timing detail, or turns a vague note into a confident medical fact. The prose may still be smooth. The bedside manner may even be excellent. Unfortunately, a hallucination delivered in fluent patient-friendly language is not safer because it has better manners. ...

June 2, 2026 · 17 min · Zelina
Cover image

Hard Problems Pay Better: Why Difficulty-Aware DPO Fixes Multimodal Hallucinations

Training data has a bad habit: the easiest examples talk the loudest. Anyone who has trained a model on preference pairs knows the scene. One answer is clearly grounded in the image; the other confidently invents an object, a color, or an action that is not there. The model learns the contrast quickly. Everyone applauds. The loss goes down. The dashboard looks obedient. ...

January 5, 2026 · 15 min · Zelina
Cover image

Planning Before Picking: When Slate Recommendation Learns to Think

A list of individually excellent items can still be a terrible list. Ask anyone who has attended a conference with five brilliant speakers, no agenda, and three consecutive sessions on the same topic. Recommendation systems have the same problem. A conventional recommender can assign highly accurate scores to individual videos, products, or articles, then still assemble a repetitive, badly ordered, or strangely balanced feed. Each item wins its private competition. The user receives the collective consequences. ...

January 2, 2026 · 18 min · Zelina
Cover image

When Safety Stops Being a Turn-Based Game

Jailbreaks are not polite enough to wait their turn. That is the awkward weakness in many safety-training pipelines. A model is attacked, patched, tested, and released. Then another attack appears, usually crafted with more creativity than the previous defense assumed. The safety team patches again. The benchmark improves. The real attack surface moves. Everyone calls this iteration, because “organized whack-a-mole with GPUs” sounds less respectable. ...

December 28, 2025 · 15 min · Zelina