Cover image

One Token, More Than One Memory

TL;DR for operators When inference compute is expensive but storing more parameters is acceptable, sparse parametric memory offers an attractive trade: increase what the model can store without activating all of that capacity for every token. The unresolved problem is retrieval quality. If memory is indexed only by token identity or a fixed local n-gram, the same surface token can repeatedly access the same stored representation even when its meaning changes with context. ...

September 30, 2026 · 8 min · Zelina
Cover image

Before the Word Arrives: How LLMs Use Sound to Choose a or an

TL;DR for operators Choosing between a and an depends on how the next word sounds, even when spelling misleads: a university but an hour. Kim and Lee find that a single sound-related direction learned from ordinary English cases generalizes to these spelling-sound exceptions, reaching 100.0% accuracy for Llama, 95.1% for Qwen, and 98.0% for Gemma.1 More importantly, this feature is not merely decodable. When researchers hold a synthetic nonce embedding fixed and change only its position along that direction, the models shift between preferring a and an. ...

September 28, 2026 · 7 min · Zelina
Cover image

The Model Felt the Tampering. It Couldn’t Name the Cause

TL;DR for operators Suppose middleware silently rewrites part of an AI agent’s own generated answer before the model continues. A reasonable expectation is that a capable model would notice the interference, or at least diagnose why its continuation has become strange. The Sleight of Word benchmark tests exactly that expectation.1 Across 19 open-weight instruction-tuned models, covert substitutions consistently change the models’ predictive distributions: post-swap surprisal and entropy rise for every model tested. But correct identification of the intervention is almost absent. No model exceeds 1.3% explicit switch awareness, and the pooled rate is below 0.1%. ...

September 25, 2026 · 7 min · Zelina
Cover image

More Data, Better Learner: What Children Reveal About Learning Efficiency

TL;DR for operators Adding more data and becoming better at learning from data are different objectives. Across five longitudinal vocabulary datasets covering American English, Norwegian, and Japanese, young children show strongly increasing returns to developmental experience. Using a matched per-word estimator, the language models examined in the study produce a median acceleration estimate of 1.16, with an interquartile range of 0.93–1.46. Corrected child estimates fall around 10.4–13.8. ...

September 24, 2026 · 7 min · Zelina
Cover image

English-Like Is Not the Same as Correct: Measuring Multilingual Reasoning Quality

TL;DR for operators A multilingual reasoning system may generate several plausible traces for the same request. The difficult part is deciding which trace to reward, trust, or return. The evidence here argues against using one English-derived trace-quality score as that decision rule. Across ten languages, many reasoning features are associated with correctness in broadly similar directions, but their effect sizes vary substantially and sometimes reverse. English semantic and structural similarity are generally positive signals, so the paper does not show that English-like reasoning is harmful. It shows that English similarity is not consistently the strongest signal. ...

September 17, 2026 · 7 min · Zelina
Cover image

When the Next Sparse Billion Should Go to Memory, Not Experts

TL;DR for operators A sparse-model team with a fixed marginal parameter budget should not assume that the next increment belongs in more experts. In the LongCat-Flash experiments reported by Liu et al.,1 parameter-equivalent expert scaling performs better earlier, but N-gram embedding scaling takes the lead after the MoE reaches a sufficiently high-sparsity regime. ...

September 11, 2026 · 8 min · Zelina
Cover image

When Conservatism Wins: Rare-Event Estimation Depends on What You Fear Missing

TL;DR for operators A model-risk team estimating failures too rare for ordinary sampling has to decide which kind of error it can least afford. The benchmark shows that this choice can reverse which estimator looks best. Under approximately symmetric penalties, GA-AMLS has much lower average SPB loss than QLD for the 1-layer and 4-layer models. When the loss instead makes underestimation much more costly, QLD sharply outperforms GA-AMLS—even though QLD systematically overestimates and GA-AMLS has lower bias. ...

August 17, 2026 · 6 min · Zelina
Cover image

Common Is Not Defining: Testing Whether Language Models Understand Category Relations

TL;DR for operators A model-review team may need to decide whether a feature is essential to a category or merely common in the data. That distinction matters because a strong association can otherwise become an unsupported ontology rule, automated policy, risk classification, or product requirement. For six embedding-based transformer models, scores initially appeared to separate defining properties from properties that were only statistically common. Once researchers controlled for human-rated prevalence—how often each property occurs—most of that separation disappeared. The same raw score that seemed to reveal conceptual structure was largely explained by frequency. GPT-4 retained a substantially stronger distinction under the same control. ...

August 8, 2026 · 6 min · Zelina
Cover image

If Logic Were Enough: Why LLMs Still Miss the Point of Conditionals

A promise is rarely just a logical operator. “If you mow the lawn, I’ll give you 50 dollars” does not sound like a philosophical exercise in truth tables. It sounds like a deal. Most people hear it as: no mowing, no money. By contrast, “If you’re hungry, there’s pizza in the oven” does not mean the pizza appears only under the metaphysical condition of your hunger. It means the pizza is there, and your hunger merely explains why I am telling you. ...

May 29, 2026 · 16 min · Zelina
Cover image

Prompt and Circumstance: Why One Accuracy Number Is Not a Reliability Audit

Opening — Why this matters now The AI market has learned to worship benchmark tables with the solemnity once reserved for quarterly earnings. One model is up two points on MMLU, another is slightly better at reasoning, a third is cheaper, smaller, faster, and therefore apparently ready to run your compliance workflow by Tuesday. ...

May 7, 2026 · 14 min · Zelina