TL;DR for operators
When inference compute is expensive but storing more parameters is acceptable, sparse parametric memory offers an attractive trade: increase what the model can store without activating all of that capacity for every token. The unresolved problem is retrieval quality. If memory is indexed only by token identity or a fixed local n-gram, the same surface token can repeatedly access the same stored representation even when its meaning changes with context.
MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup1 addresses this by keeping the cheap first-stage lookup and making only the second stage contextual. Each indexed memory row contains several candidate vectors. The model’s current hidden state chooses a small subset, combines them, and separately controls how strongly that retrieved memory modifies each attention value head.
The controlled evidence is encouraging within the tested regime. With 151M memory parameters on the nanochat-style backbone, Bigram reached 0.8636 validation bits per byte and Value Embedding 0.8633. MoME reached 0.8621 with standard token rows and 0.8611 with its $c_{\mathrm{grp}}=2$ grouping configuration. That grouped configuration also produced the highest CORE score among the matched-memory rows, 0.1583 versus 0.1533 for Bigram and 0.1522 for Value Embedding.
For product teams, the relevant question is not whether sparse memory is generally good. It is whether adding context-sensitive stored capacity improves model quality enough to justify additional storage, training complexity, and latency. This paper provides evidence that the trade can work at sub-billion scale and in one approximately 100B-token training run. It does not establish that the same scaling relationship will hold at frontier scale.
Sparse capacity still needs the right memory for the context
Sparse memory changes the economics of model capacity because stored parameters do not all need to participate in every forward pass. A model can therefore carry substantially more learned information than its active computation would suggest.
The limitation is that cheap addressing is usually coarse. A token-indexed table retrieves according to token identity. An n-gram table adds some local context, but its address is still determined before the model considers the broader hidden representation.
That creates a concrete representation problem. A token such as bank can appear in a financial context or beside a river. If both occurrences resolve to one fixed token memory, added storage capacity does not by itself provide context-specific retrieval.
MoME changes the organization inside the memory row rather than discarding inexpensive indexing.
MoME keeps token lookup deterministic and makes slot selection contextual
MoME’s memory is structured as
so each of the $N$ addressable rows contains $M$ candidate memory vectors instead of one.
The first lookup remains inexpensive and deterministic. In the canonical design, the current token selects a row. Context enters after that lookup: for each value head, a learned router reads the current hidden state $\mathbf{h}_t$, scores the slots in the selected row, and activates only the top $K$.
The activated memories are then combined into a context-specific vector. A second learned mechanism controls injection strength:
Retrieval and use are therefore separate decisions. The router determines which stored vectors are relevant; the residual gate determines how strongly that retrieved information should alter the value stream.
This is also why MoME should not be read as an internal version of retrieval-augmented generation. It does not search an external corpus, and it does not replace token addressing with unrestricted semantic retrieval. The memory vectors are learned parameters, and context selects among candidates only after the inexpensive first-stage row has already been identified.
The controlled gains survive more than one backbone
The nanochat comparison provides the cleanest matched-memory result.
| Method | Memory parameters | Validation bpb | CORE | Training throughput |
|---|---|---|---|---|
| Bigram | 151M | 0.8636 | 0.1533 | 1.09M tok/s |
| Value Embedding | 151M | 0.8633 | 0.1522 | 1.10M tok/s |
| MoME, $c_{\mathrm{grp}}=1$ | 151M | 0.8621 | 0.1571 | 1.07M tok/s |
| MoME, $c_{\mathrm{grp}}=2$ | 151M | 0.8611 | 0.1583 | 1.07M tok/s |
Lower bpb is better. At the same 151M memory budget, the MoME variants improve validation bpb over both deterministic baselines. The strongest grouped configuration also improves aggregate CORE.
Increasing MoME memory to 302M parameters reduces validation bpb further. The $c_{\mathrm{grp}}=2$ configuration reaches 0.8561 validation bpb and 0.1664 CORE, the strongest reported combination among the 302M MoME grouping variants.
The result is not confined to that backbone. Across Llama/MobileLLM 125M, MobileLLM 350M, and Qwen3 0.6B experiments, MoME improves validation bpb and CORE over STEM while using fewer memory parameters. Comparisons with Value Embedding are less uniform across metrics and configurations, so the evidence supports a general advantage over several deterministic alternatives rather than universal dominance in every matched comparison.
A separate nanochat memory-size sweep also finds lower validation bpb for MoME than matched-memory Bigram at every tested memory capacity. Importantly, Bigram and MoME are not mutually exclusive: the paper also uses Bigram as the first-stage indexer and then applies contextual slot routing inside the retrieved row. That experiment supports the architectural claim that n-gram addressing and contextual memory selection can be composed.
The larger training run tests scale, not a universal scaling law
The approximately 100B-token experiment is a useful extension because it tests the mechanism under substantially longer training.
With the same roughly 0.78B dense network, Value Embedding adds 0.60B memory parameters and MoME 0.61B. MoME records lower bpb than both Dense and Value Embedding on FineWeb-Edu, enwik9, and Shakespeare, and its CORE-22 score reaches 0.3687 versus 0.3484 for Value Embedding and 0.3441 for Dense.
The comparison is not one-sided. Value Embedding records slightly lower in-domain ClimbMix bpb: 0.6463 versus 0.6504 for MoME.
This experiment should therefore be interpreted as a scale check supporting the smaller controlled results, not proof of an invariant scaling advantage. It uses a single seed, and the underlying dense backbone remains about 0.8B parameters.
Context sensitivity appears in the router, but causality is unresolved
The paper also asks whether the router is merely distributing traffic among extra parameters or whether its choices contain semantic structure.
Qualitative examples for words including bank, bug, drive, apple, cell, and python show different slots becoming active under different contextual uses.
The quantitative Word-in-Context analysis is more informative. After strict filtering, it retains 670 sentence pairs. Different-sense pairs exhibit greater routing-distribution divergence at 117 of 144 layer-head sites. Same-sense pairs exhibit greater chance-corrected top-2 slot overlap at 110 of 144 sites.
Those patterns are consistent with the intended mechanism: hidden context changes which memory slots are used.
They do not establish that these routing differences cause the model’s performance gains. The probe is observational, combines WiC train and test examples, is heavily noun-skewed, and examines many layer-head sites without a held-out confirmatory analysis. The evidence supports semantic sensitivity in routing, not a causal account of why MoME predicts better.
For deployment, the trade is storage for conditional capacity
Cognaptus’ interpretation is that MoME is most relevant where storage and active compute have materially different costs.
A team deploying a compact or on-device model may be able to accommodate a larger parameter store while remaining tightly constrained by per-token computation. In that environment, increasing dense width or depth activates more computation continuously. Sparse memory instead offers additional learned capacity that is touched selectively.
The runtime measurements make this more than a parameter-count argument. In the principal nanochat comparison, MoME runs at 1.07M tokens per second versus 1.09M for Bigram with the same 151M memory budget, a difference of roughly 2%. The paper’s inference measurements also show relative overhead declining on larger tested backbones: value-stream injection adds 7.82% latency on the 350M Llama/MobileLLM case, 2.35% on Qwen3 4B, and 1.56% on Qwen3 8B.
These figures do not guarantee favorable economics on a specific serving stack. They indicate that the architecture can add conditional stored capacity without turning every added parameter into proportional active computation, and that implementation placement can materially affect whether the resulting latency is acceptable.
The evidence stops before frontier-scale deployment
Three boundaries should constrain adoption decisions.
First, most controlled experiments use dense backbones below one billion parameters. The approximately 100B-token experiment extends training scale, but not dense-backbone scale by an equivalent amount.
Second, the strongest scaling evidence covers only the memory capacities actually tested. A better slope than Bigram across those points should not be converted into a general scaling law.
Third, semantic routing is evidence about mechanism behavior, not yet a dependable control interface. The paper shows that contextual sense and routing patterns are associated. It does not demonstrate that a product team can inspect, edit, or steer those slots and reliably obtain predictable behavioral changes.
Within those limits, MoME sharpens the design question around sparse parametric memory. The choice is no longer only how much memory to add or how cheaply to index it. It is also whether the same inexpensive address should expose multiple context-dependent memories. In the tested regimes, making that second decision conditional improves the quality-capacity tradeoff without eliminating the efficiency advantage that made sparse memory attractive in the first place.
Cognaptus: Automate the Present, Incubate the Future.
-
Muchen Li and Leonid Sigal and Renjie Liao (2026). MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup. arXiv:2609.15126. https://arxiv.org/abs/2609.15126 ↩︎