Cover image

Teaching the Query When to Move: Small Models as Retrieval Controllers

TL;DR for operators A retrieval agent does not merely retrieve documents. It repeatedly decides what to do next: search again, rewrite a query, record evidence, combine what it has found, verify a claim, or stop. The Fellowship of the Query: Learning Retrieval Actions1 shows that these decisions can be taught to small language models as a supervised classification problem. Across the evaluated models, LoRA fine-tuning sharply improves prediction of the next retrieval action. For Granite 4.1 3B, macro-F1 rises from 0.1736 zero-shot to 0.6536 after fine-tuning. ...

October 3, 2026 · 7 min · Zelina
Cover image

Probe Before You Prune: Measuring Which MoE Experts a Workload Can Lose

TL;DR for operators A mixture-of-experts model can be too large for an inference budget even though only a small subset of its experts executes for each token. The inactive experts still have to reside somewhere. In the measured Mixtral deployment studied by Janati et al., removing four of eight experts per layer cut 4-bit memory from 24.2 GB to 12.3 GB and per-token latency from 40.3 ms to 25.2 ms on a single A100.1 ...

September 6, 2026 · 9 min · Zelina
Cover image

The Optimizer Comes Before the Quantizer

TL;DR for operators A small model can perform acceptably after task specialization and still lose more accuracy than expected when compressed. The paper examined here suggests that part of that loss may be shaped earlier, during fine-tuning, rather than being determined by the quantizer alone. Paper evidence: Muon-optimized models lost less accuracy after quantization than Adam-optimized models on six of eight benchmarks. On ARC-e, for example, accuracy fell by 0.55 percentage points after quantization with Muon versus 3.16 points with Adam. The integrated pipeline also produced a 2.86 GB quantized model from a 6.01 GB pre-quantized model while improving reported serving throughput and roughly halving per-token latency. ...

September 6, 2026 · 8 min · Zelina
Cover image

When Confidence Drives the Workflow: What HypeLoRA Changes About Adapter Selection

TL;DR for operators A classifier can score better on a benchmark while giving probabilities that systematically overstate or understate how often it is correct. That matters whenever confidence feeds an escalation threshold, automatic action, human-review queue, or other operational rule. LoRA adapts a frozen model through a small update built from two low-rank factors, $A$ and $B$. HypeLoRA asks whether generating those adaptations across layers with a shared hyper-network improves not just task performance but the reliability of the resulting confidence signal.1 The answer is conditional: standard LoRA does not calibrate uniformly better than full fine-tuning, and fully generated HypeLoRA remains broadly similar to ordinary LoRA. The strongest reported calibration result comes from the Transformer fixed-$A$ configuration, which freezes $A$ and generates only $B$: ECE reaches 0.100 on CoLA and 0.028 on SST-2, while task performance remains below the strongest-performing configuration. ...

September 6, 2026 · 7 min · Zelina
Cover image

Compress the Activation, Spend the Memory on Rank

TL;DR for operators A LoRA fine-tuning job can fit its trainable parameters comfortably on a GPU and still run out of memory because backpropagation retains large intermediate activations. CARE-LoRA attacks that remaining buffer rather than shrinking the adapter itself. Zhang et al. store LoRA’s already-compressed activation plus a small reconstruction matrix, then use those tensors to approximate only the gradient needed for one of LoRA’s two trainable matrices.1 ...

September 5, 2026 · 7 min · Zelina
Cover image

LoRA’s Missing Budget: Which Matrices Deserve an Adapter?

TL;DR for operators LoRA already avoids the cost of updating an entire pretrained model, but it can still spend adapter capacity uniformly across matrices that do not appear equally responsive to low-rank changes. If a training team has a fixed fine-tuning budget, there are therefore two allocation decisions: how large each adapter should be, and which matrices should receive one. ...

September 5, 2026 · 7 min · Zelina
Cover image

Train Wide, Deploy Narrow: LoRA Rank Does Not Have to Be One Decision

TL;DR for operators A production team may want every LoRA adapter to fit a small, uniform serving footprint. The usual response is to choose that small rank before training and optimize inside the resulting constraint. This paper shows that the training capacity and the deployment capacity do not always need to be identical. ...

September 5, 2026 · 7 min · Zelina
Cover image

Privacy Starts Before the First Gradient

TL;DR for operators Federated fine-tuning keeps raw examples on the client, but that does not mean the client begins from a neutral model state. A malicious coordinating server can send an adapter deliberately structured so that private examples produce recoverable traces during training. Privacy risk can therefore enter through what the client downloads, not only through what it later uploads. ...

August 28, 2026 · 6 min · Zelina
Cover image

Stored Is Not Reachable: Why Continual Fact Writing Breaks Under Later Updates

TL;DR for operators A product team may repeatedly fine-tune a deployed model with new policies, prices, procedures, or customer facts, then approve each update because the newest item can be recalled immediately. That test shows the write changed current behavior. It does not show that earlier facts will remain usable after later updates. ...

August 2, 2026 · 10 min · Zelina
Cover image

The Smart Part Was the Memory, Not the Controller

TL;DR for operators A deployed model encounters a familiar operating condition again. Should it relearn the task, search for a saved configuration, or let a trained controller decide which parameters to reuse? The paper’s strongest result points to the simpler mechanism. Removing the searchable store of compressed task-specific configurations increased recovery from 1.27 to 13.33 adaptation steps—close to the 14.27 steps required by a baseline without that store. The component that looks least intelligent therefore accounts for most of the recovery advantage: retaining the small parameter modules that worked before and restoring them when the task returns. The paper calls this store the TaskKnowledgeBank. ...

August 1, 2026 · 9 min · Zelina