Cover image

One Token, More Than One Memory

TL;DR for operators When inference compute is expensive but storing more parameters is acceptable, sparse parametric memory offers an attractive trade: increase what the model can store without activating all of that capacity for every token. The unresolved problem is retrieval quality. If memory is indexed only by token identity or a fixed local n-gram, the same surface token can repeatedly access the same stored representation even when its meaning changes with context. ...

September 30, 2026 · 8 min · Zelina