Where to Go Deeper Beyond This Academy

A curated guide to textbooks, authors, websites, and papers for readers who want to study transformer internals, attention math, fine-tuning, GPU optimization, and benchmarking in more depth.

April 23, 2026 · 9 min · Michelle
Cover image

Where You Pause Changes What You Forget

TL;DR for operators Kim and colleagues’ study of pause-token fine-tuning1 points to a training rule that is easy to miss if pause tokens are treated mainly as extra thinking time. At an equal pause-token budget, putting pauses at semantic boundaries is the only tested placement that consistently improves both math and code averages over ordinary supervised fine-tuning. The stronger variant also masks the loss on the pause tokens themselves. ...

October 1, 2026 · 6 min · Zelina
Cover image

When the Proxy Wins: Attention Sensitivity Can Hit the Ceiling and Still Miss ICL

TL;DR for operators An internal metric can track a capability well enough to be useful for diagnosis and still fail once the training loop is told to maximize it. In this study, an attention-based signal intended to reflect whether demonstrations matter to a model rises from about 0.516 to 1.413 under direct optimization—within 0.5% of its theoretical maximum. Yet the corresponding behavioural effect is essentially absent on the controlled QA probe, and MMLU accuracy falls to 0.279.1 ...

September 23, 2026 · 7 min · Zelina
Cover image

The Fourth Hop Changes the Risk Profile: Measuring Reliability in Multi-Step LLM Workflows

TL;DR for operators When one model output becomes input to the next stage, a final accuracy score tells you too little about where reliability is being lost. A workflow may fail because a required fact was never available, because a later composition step is intrinsically harder, or because an earlier mistake was allowed to propagate. Those failure modes call for different controls. ...

September 17, 2026 · 8 min · Zelina
Cover image

Fine-Tuning Changes What Your Model’s Errors Reveal

TL;DR for operators A fine-tuned model can become only slightly more accurate while its remaining errors become substantially easier to distinguish from correct answers. That matters when uncertainty scores feed operational controls. If a production workflow accepts an answer, abstains, calls another model, or sends a case to human review according to a detector threshold, fine-tuning changes more than the benchmark score. It can change the detector itself as an operating signal. ...

September 10, 2026 · 7 min · Zelina
Cover image

Compress the Activation, Spend the Memory on Rank

TL;DR for operators A LoRA fine-tuning job can fit its trainable parameters comfortably on a GPU and still run out of memory because backpropagation retains large intermediate activations. CARE-LoRA attacks that remaining buffer rather than shrinking the adapter itself. Zhang et al. store LoRA’s already-compressed activation plus a small reconstruction matrix, then use those tensors to approximate only the gradient needed for one of LoRA’s two trainable matrices.1 ...

September 5, 2026 · 7 min · Zelina
Cover image

LoRA’s Missing Budget: Which Matrices Deserve an Adapter?

TL;DR for operators LoRA already avoids the cost of updating an entire pretrained model, but it can still spend adapter capacity uniformly across matrices that do not appear equally responsive to low-rank changes. If a training team has a fixed fine-tuning budget, there are therefore two allocation decisions: how large each adapter should be, and which matrices should receive one. ...

September 5, 2026 · 7 min · Zelina
Cover image

Train Wide, Deploy Narrow: LoRA Rank Does Not Have to Be One Decision

TL;DR for operators A production team may want every LoRA adapter to fit a small, uniform serving footprint. The usual response is to choose that small rank before training and optimize inside the resulting constraint. This paper shows that the training capacity and the deployment capacity do not always need to be identical. ...

September 5, 2026 · 7 min · Zelina
Cover image

Synthetic Data Can Make the Model Worse

TL;DR for operators A team with authoritative domain documents but little labeled training data has an attractive option: ask a capable model to manufacture question-answer pairs, then fine-tune a smaller open model on them. The operational risk is assuming that domain relevance makes those examples safe training material. In this paper, a simple synthetic-data pipeline moved LLaMA 3.1 8B backward on open-ended legal QA: its LegalMC4 score fell from 43.0% to 35.4%. A more structured pipeline raised the same score to 55.4%. Across both LLaMA 3.1 8B and Gemma 3 12B, that structured treatment improved all four tested German legal benchmarks.1 ...

September 3, 2026 · 7 min · Zelina
Cover image

One Correction, Every Case: When LLMs Actually Update the Rule

TL;DR for operators An AI system receives one signal that an operating rule has changed. The important test is not whether its average performance eventually recovers, but whether it immediately applies the revised rule to cases it has not yet revisited. Many models fail this test quietly. They correct each stimulus only after encountering it again, producing gradual recovery without inferring that one hidden rule changed for every stimulus at once. For teams deploying agents, that distinction matters whenever a policy change, workflow update, or exception rule must propagate across related cases. ...

July 31, 2026 · 10 min · Zelina