Cover image

Teach the Restoration, Skip the Runtime Step

TL;DR for operators ReCAST1 treats de-obfuscation as something a classifier should learn during training, not necessarily something a production system should execute before every classification. A Qwen3.5-9B student first learns to identify disguised spans, classify the obfuscation, reconstruct normalized text, and assign a risk label. It is then adapted to classify the original obfuscated SMS directly. ...

September 25, 2026 · 7 min · Zelina
Cover image

Guide, Don’t Guess: Reallocating Reasoning Compute with Small-Model Hints

TL;DR for operators When a compact model fails on a difficult multi-step problem, the default remedies are expensive: use a larger model, generate many independent attempts, or accept lower accuracy. This paper tests a different allocation of compute—keep the solver small, but give it localized guidance at difficult intermediate steps. HintMR1 separates those jobs. A hinter tells the solver what to consider next without supplying the full solution; the solver then advances its reasoning one step at a time. On AIME-2024, DeepSeek-R1-Distill-Qwen-7B rises from 20.69% accuracy without hints to 68.97% with GPT-5.2-generated hints. The result is not evidence that any second model helps: non-fine-tuned small-model hints are inconsistent and sometimes reduce accuracy below the no-hint baseline. ...

September 18, 2026 · 6 min · Zelina
Cover image

Train the Family, Not Every Size From Scratch

TL;DR for operators A team that needs small, medium, and large versions of the same foundation model would normally budget several substantial pre-training runs. That accounting assumes each model must relearn most of the family’s shared representation from scratch. Chain-of-Models Pre-Training (CoM-PT)1 challenges that assumption. The smallest model is trained normally; each larger model then inherits parameters and feature guidance from the immediately smaller predecessor. On CC3M, the paper’s four-model ViT family cuts accumulated training MACs from 16.70 to 2.94 in the reported units, equivalent to 5.68x accumulated acceleration. Its ViT-L/16 endpoint also exceeds the individually trained baseline on ImageNet-1K, VTAB+, and COCO. ...

September 12, 2026 · 8 min · Zelina
Cover image

The Optimizer Comes Before the Quantizer

TL;DR for operators A small model can perform acceptably after task specialization and still lose more accuracy than expected when compressed. The paper examined here suggests that part of that loss may be shaped earlier, during fine-tuning, rather than being determined by the quantizer alone. Paper evidence: Muon-optimized models lost less accuracy after quantization than Adam-optimized models on six of eight benchmarks. On ARC-e, for example, accuracy fell by 0.55 percentage points after quantization with Muon versus 3.16 points with Adam. The integrated pipeline also produced a 2.86 GB quantized model from a 6.01 GB pre-quantized model while improving reported serving throughput and roughly halving per-token latency. ...

September 6, 2026 · 8 min · Zelina
Cover image

You Can’t Reweight a Dead End: TRD and the Prefix Failure Problem

TL;DR for operators The paper’s main message is simple: if a reasoning model has already walked into a dead end, per-token distillation often keeps supervising it from inside the dead end. A clever loss cap is not a map. A top-k filter is not a tow truck. Trajectory-Refined Distillation, or TRD, repairs the student’s own rollout before using it for distillation. The pipeline is: sample the student’s attempt, ask a teacher or privileged self-teacher to rewrite the trajectory into a better one, then train on the refined trajectory rather than on the original failed rollout. The technical contribution is not “better prompting”, although prompts are used. It is the shift from token-level correction to trajectory-level correction. ...

June 19, 2026 · 15 min · Zelina
Cover image

The Rule Is the Model: DEM’s Case for Bedside Anomaly Detection Without Explainer Theatre

Alerts are cheap; trusted alerts are not A hospital monitor that screams without explaining itself is not a decision-support system. It is a very expensive doorbell. That is the practical problem behind Singh, Roy, Bose, and Hota’s Distilled Explanation Model, or DEM, for physiological anomaly detection in wireless body area networks.1 The paper is nominally about clinical sensor data: heart rate, oxygen saturation, blood pressure, temperature, stress signals, sensor dropouts, and ICU monitoring. But the more interesting argument is architectural. DEM is not trying to make a black-box model more charming after it has already made a decision. It is trying to make the explanation part of the decision itself. ...

June 14, 2026 · 17 min · Zelina
Cover image

When 256 Dimensions Pretend to Be 16: The Quiet Overengineering of Vision-Language Segmentation

A prompt is usually a small thing. “White dog.” “Person in a blue jacket.” “Cup on the table.” Nobody hears these phrases and thinks: excellent, time to deploy a large general-purpose language encoder. Yet that is often what modern vision-language segmentation systems do. The visual model may be carefully optimized. The deployment team may obsess over image encoder latency, GPU memory, and batch size. Then the text side sits there, inherited from a larger foundation model stack, quietly burning capacity to understand what is often a noun phrase with a color adjective attached. Very sophisticated machinery, bravely parsing “red car.” Heroic. ...

February 13, 2026 · 15 min · Zelina
Cover image

When Models Listen but Stop Thinking: Teaching Audio Models to Reason Like They Read

A voice assistant can transcribe your question correctly and still answer like it heard something else. That is the awkward part of modern audio-language models. The obvious diagnosis is usually “better speech recognition.” The less obvious diagnosis is nastier: the model may receive an audio input that is semantically equivalent to the text prompt, but once generation begins, its audio-conditioned reasoning trajectory drifts away from the reasoning trajectory it would have followed if the same question had been typed. ...

January 26, 2026 · 16 min · Zelina