Cover image

Grammar Before Language: What Multilingual LLMs Decide First

TL;DR for operators A translation can be wrong in at least three different ways: the model can arrange words incorrectly, produce the wrong language, or choose the wrong lexical content. The paper examined here finds evidence that these are not merely different visible error types. In the tested multilingual LLMs, they often correspond to separable internal stages. ...

September 30, 2026 · 7 min · Zelina
Cover image

Same Proposition, Different Stance: Grammar as a Model-Risk Variable

TL;DR for operators Huang, Padó, and Weeber show that changing the grammatical realization of a political statement can change an LLM’s expressed stance even when the rewrite is intended to preserve the proposition.1 Across four Gemma-3 and Qwen3 models, active/passive conversions and cleft constructions produced measurable stance movement, and variance decomposition indicates that the effect is systematic rather than ordinary run-to-run sampling noise. ...

August 20, 2026 · 7 min · Zelina
Cover image

Heads You Lose: Why Ablation-Reversible Interpretability Doesn’t Transfer

TL;DR for operators The paper is a useful slap on the wrist for anyone tempted to turn an interpretability result into an operational control too quickly.1 It asks a simple question: when an attention head looks important, contains readable information, and can restore model behaviour after ablation, does that mean it carries a transferable representation of the computation? ...

June 17, 2026 · 17 min · Zelina
Cover image

How Sparse is Your Thought? Cracking the Inner Logic of Chain-of-Thought Prompts

TL;DR for operators Chain-of-thought prompting is often sold as a window into model reasoning. This paper is more useful because it treats CoT as something less mystical and more testable: a prompt-induced change in internal representations.1 The researchers train sparse autoencoders on hidden activations from two Pythia models solving GSM8K math problems under CoT and NoCoT prompts. They then patch CoT-derived sparse features into NoCoT runs and ask a sharper question: does inserting those internal features increase the log-probability of the correct answer? ...

August 1, 2025 · 16 min · Zelina