Cover image

Merge Without Mayhem: How Orthogonal Deltas Could Revolutionize Model Composition

TL;DR for operators Model composition usually sounds harmless until someone asks the obvious production question: “Can we remove that client-specific update without retraining the whole thing?” At that point, many elegant AI stacks quietly become sedimentary rock. The MDM-OC paper proposes a cleaner model lifecycle: keep a shared base model, express every fine-tuned specialist as a task delta, orthogonalize those deltas so they interfere less, merge them with tuned coefficients, and subtract a selected delta later when a capability, customer, or data source needs to be removed.1 The important claim is not “we found another averaging recipe.” The claim is that model updates can be treated as separable components in parameter space. ...

August 2, 2025 · 20 min · Zelina
Cover image

Reasoning at Scale: How DeepSeek Redefines the LLM Playbook

TL;DR for operators DeepSeek-R1 is not a story about one model suddenly becoming clever because someone found the secret lever labelled “reason harder”. It is a systems story: take a strong base model, reward it on problems where correctness can be checked, let longer reasoning traces emerge, repair the ugly parts with cold-start data and alignment, then distil the resulting behaviour into smaller models where deployment economics actually matter.1 ...

July 15, 2025 · 14 min · Zelina
Cover image

The Outlier Is a Lie: Quantization Breakthroughs with OSP

TL;DR for operators If your deployment plan depends on squeezing a language model into cheap inference hardware, this paper is worth reading because it changes the timing of the quantization problem. Most quantization work asks: “How do we repair a model after training so it survives 4-bit inference?” Outlier-Safe Pre-Training asks a more irritating question: “Why did we train a quantization-hostile model in the first place?”1 ...

June 25, 2025 · 18 min · Zelina