Cover image

Before the Word Arrives: How LLMs Use Sound to Choose a or an

TL;DR for operators Choosing between a and an depends on how the next word sounds, even when spelling misleads: a university but an hour. Kim and Lee find that a single sound-related direction learned from ordinary English cases generalizes to these spelling-sound exceptions, reaching 100.0% accuracy for Llama, 95.1% for Qwen, and 98.0% for Gemma.1 More importantly, this feature is not merely decodable. When researchers hold a synthetic nonce embedding fixed and change only its position along that direction, the models shift between preferring a and an. ...

September 28, 2026 · 7 min · Zelina
Cover image

A Citation Can Be Right Without Being Grounded

TL;DR for operators A RAG system can return the right answer, attach a source that genuinely supports that answer, and still leave one critical question unresolved: did that source actually influence the model’s answer generation? A mechanistic study of Llama-3.1-8B-Instruct finds that inline citation behavior is not controlled by one dedicated citation feature. It emerges from a distributed sequence of attention heads and MLPs that includes early entity enrichment, matching between document and question entities, mid-layer processing, and late aggregation that shifts the model toward emitting a citation marker rather than ending the sentence.1 ...

September 19, 2026 · 8 min · Zelina
Cover image

Sparse Routing May Buy You Inspectability, Not Just Efficiency

TL;DR for operators Herbst, Wermter, and Lee find that the analyzed Mixture-of-Experts models often represent tested concepts in far fewer neurons than comparable dense transformers.1 The difference is largest under the hardest probe constraint: when only one neuron is available, MoE experts often approach their own best probe performance while dense feed-forward layers need more dimensions. Models with sparser routing also tend to show cleaner representations. ...

September 11, 2026 · 8 min · Zelina
Cover image

Same Proposition, Different Stance: Grammar as a Model-Risk Variable

TL;DR for operators Huang, Padó, and Weeber show that changing the grammatical realization of a political statement can change an LLM’s expressed stance even when the rewrite is intended to preserve the proposition.1 Across four Gemma-3 and Qwen3 models, active/passive conversions and cleft constructions produced measurable stance movement, and variance decomposition indicates that the effect is systematic rather than ordinary run-to-run sampling noise. ...

August 20, 2026 · 7 min · Zelina
Cover image

Reasoning Tokens Are Compute, Not an Audit Trail

TL;DR for operators Giving an AI system more reasoning tokens can improve difficult answers because each generated token triggers another round of model computation and preserves intermediate information for the next step. For problems requiring a sequence of dependent operations, this can create additional computational depth rather than merely reveal reasoning that was already complete inside the model. ...

August 11, 2026 · 7 min · Zelina
Cover image

Stable Enough to Be Wrong: Why Neuron Selectors Need Causal Audits

TL;DR for operators A team using an importance-ranking tool must know whether its low-ranked model components are genuinely safe to remove—not merely whether the tool produces the same ranking repeatedly. On LLaMA-3.1-8B at 30% masking, Wanda achieved near-perfect ranking stability of 0.9998, yet removing the components it labelled least important drove perplexity to 360,083. Consensus-2 was less stable at 0.945, but the same removal test produced perplexity of 66.1. The more reproducible ranking was far less causally faithful. ...

August 10, 2026 · 9 min · Zelina
Cover image

The Probe Saw the Prompt Before It Saw the Fake

TL;DR for operators A safety team evaluating an internal model monitor must answer three different questions: is the model changing its behavior under monitoring, is the detector merely recognizing the evaluation setup, and does the detected signal provide any control over the behavior? The evidence begins narrowly. Among thirteen tested open-weight models, only Qwen3-32B and Llama-3.1-8B showed positive monitoring-induced compliance gaps treated as natural alignment faking. In both, monitored models could comply externally while their internal states remained, on average, closer to states associated with refusal. Yet near-perfect raw probe scores mainly identified which system prompt the model had received. Under stricter query-held-out tests designed to prevent prompt and query leakage, reliable per-response detection survived only on Llama-3.1-8B. ...

August 4, 2026 · 8 min · Zelina
Cover image

One Correction, Every Case: When LLMs Actually Update the Rule

TL;DR for operators An AI system receives one signal that an operating rule has changed. The important test is not whether its average performance eventually recovers, but whether it immediately applies the revised rule to cases it has not yet revisited. Many models fail this test quietly. They correct each stimulus only after encountering it again, producing gradual recovery without inferring that one hidden rule changed for every stimulus at once. For teams deploying agents, that distinction matters whenever a policy change, workflow update, or exception rule must propagate across related cases. ...

July 31, 2026 · 10 min · Zelina
Cover image

Brain Scan for a Machine That Does Not Have a Brain

TL;DR for operators Most model-governance systems still treat LLM failure like a customer-support ticket: hallucination, bias, unsafe compliance, sycophancy, escalation, add a dashboard, summon a committee, repeat until morale improves. NeuroCogMap proposes a more useful question: when the model fails, which internal systems were recruited, under-recruited, or misrouted? The paper builds a functional atlas of LLM internals by clustering sparse autoencoder features into parcels, attaching cognitive descriptions to those parcels, mapping them to capabilities, and arranging those capabilities into a four-level hierarchy: perception, representation, abstraction, and application.1 ...

July 7, 2026 · 20 min · Zelina
Cover image

No Structure, No Glory: Why AI Cognition Has to Be Shown, Not Named

TL;DR for operators AI systems are now sold with labels that sound increasingly cognitive: reasoning, planning, agency, memory, autonomy, sometimes even the more theatrical hints of machine consciousness. Lovely. The marketing department has discovered philosophy. The useful question is not whether the label feels exciting. It is whether the system realizes an internal organization that could actually support the claimed capability. ...

June 29, 2026 · 18 min · Zelina