TL;DR for operators
An agent can have a compressed memory that answers the current question perfectly and is already inadequate for the next update.
Guangzhe Zhang’s Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression1 demonstrates the problem with paired synthetic histories. On DeepSeek, the paper’s frontier selector and a latest-only variant both score 96/96 on current reads. After a later update makes older evidence relevant, frontier remains at 96/96, while latest-only falls to 32/96.
For persistent assistants, this changes what memory QA needs to certify. Tests should include later revocations, rollbacks, expiries, dependency changes, and references to previously unimportant records. They should also report whether failures came from missing information, failed memory writes, or an answer-format contract. Those are different engineering problems.
The study is a controlled development experiment with 24 synthetic source-history pairs, not evidence that its selector is generally superior on natural conversations, repositories, or long-running production agents.
A correct checkpoint can hide an already-broken memory
Consider an assistant maintaining a changing operational record. It compresses the history, answers today’s query correctly, and keeps only the compact state for tomorrow. A later correction arrives. Or a revocation is replayed. Or a reference suddenly points to an older version.
At the original checkpoint, there may have been no visible reason to distinguish a good compressed state from a bad one. Both produced the same answer.
That is the paper’s central evaluation problem. It constructs two different histories that require the same answer now but contain different latent evidence. These are matched-current histories. Both then receive the same future update. A reveal future makes the hidden historical difference relevant, so the correct later answers diverge. An override future instead installs a common newer answer, checking that memory can also accept legitimate replacement.
This design asks a harder question than “Can the memory answer now?” It asks whether the memory has preserved the distinctions needed to evolve correctly.
The 96/96 tie is the result that changes the benchmark
The strongest illustration comes from the latest-only ablation, which discards older live versions.
On DeepSeek:
| Memory condition | Current | Reveal | Joint |
|---|---|---|---|
| Frontier | 96/96 | 96/96 | 45/48 |
| Latest-only | 96/96 | 32/96 | 15/48 |
Current accuracy alone declares a tie. The later update exposes a 64-case difference in reveal performance.
GLM makes the same point from another direction. Latest-only actually scores 84/96 on current reads, above frontier’s 78/96. After the reveal update, the ordering reverses sharply: 23/96 for latest-only versus 82/96 for frontier.
The paper’s mechanism-level results show that this is not generic degradation. Latest-only fails systematically where older versions matter: rollback, expiry, tombstone replay, and the mixed mechanism. Similarly, removing tombstones creates exactly 16/96 wrong retained logs on reveal, all concentrated in the tombstone-replay cases; strict reveal success there is 0/16 on both backends.
These are ablations, so their role is diagnostic. They identify which historical structures must survive compression under the paper’s event semantics. They do not establish that every production memory system should store every old record indefinitely.
A failed answer does not always mean failed memory
End-to-end evaluation creates another problem: several layers can fail after compression.
The paper therefore retrospectively separates retained-state adequacy from writer delivery and reader-schema compliance. That distinction materially changes how some scores should be interpreted.
For the GLM frontier condition, strict reveal accuracy is 82/96. The remaining 14 cases initially look like memory failures. But the diagnostic interpreter finds the correct bare values in all 14. The reader omitted the required answer wrapper. Under the alternative bare-map rule, reveal performance becomes 96/96.
The structured model-written memory has a different failure profile. Its reveal records include 26 valid but wrong logs on DeepSeek and 25 on GLM, alongside unparsed outputs and other pipeline effects. Writer diagnostics also show empty and incomplete generations.
This means the headline comparison between a deterministic selector and a model-generated writer is an end-to-end pipeline comparison. It combines content selection, generation reliability, cap compliance, delivery, and answer formatting. It cannot isolate a pure causal advantage in semantic compression.
For an engineering team, the reporting implication is straightforward: keep at least three failure classes separate.
- Retained-state failure: necessary information is no longer present or has the wrong state.
- Delivery failure: the writer never produces an acceptable memory artifact.
- Response-contract failure: the needed value exists, but the reader violates the required output schema.
Without that split, the wrong subsystem gets repaired.
Renaming identifiers reveals a shortcut that normal cases miss
The paper then tests a different question: does a high-scoring deterministic selector depend on arbitrary properties of the benchmark representation?
Under the original late-reference cases, frontier retains the required state in 8/8 variants. The authors consistently rename identifiers while preserving their semantics and serialized lengths. Adequacy falls to 94/320.
The other five tested mechanisms remain at 320/320 under the same renaming procedure. The failure is therefore localized to late-reference selection rather than evidence that renaming generally breaks the event system.
The cause is lexical identifier tie-breaking. When memory capacity forces a selection, arbitrary names can influence which records survive.
A repaired selector, frontier-alpha, removes this label dependence and records zero violations across 13,440 tested stage-level equivariance comparisons. But this repair does not solve the substantive problem. Frontier-alpha retains only 2/8 required late-reference values on the original cases, 80/320 under renaming, and 78/320 under prefix permutation.
That negative result is valuable. Robustness to irrelevant naming and sufficiency for unknown future queries are separate properties. Removing a shortcut does not create information that bounded memory failed to retain.
What changes for agent-memory QA
Cognaptus infers three practical changes from the paper’s evaluation design.
First, deployment tests for persistent agents should contain state transitions, not only static questions. If an assistant operates over mutable policies, tickets, customer records, software state, permissions, or research histories, the test suite should deliberately include later revocations, rollbacks, expiries, dependency changes, and references to older information.
Second, memory acceptance criteria should be layered. A semantic retained-state test belongs beside delivery reliability and response-contract compliance. This gives operators a path from failure to remediation rather than one composite score.
Third, deterministic memory policies deserve semantics-preserving perturbation tests. Identifier renaming, record-order changes, and related transformations can reveal dependence on conventions that carry no intended meaning. Passing such tests is not proof of future sufficiency, but failing them identifies avoidable implementation fragility.
These practices are especially relevant when comparing vendors or architectures. If two systems differ in retry policies, writer generation, access to raw archives, serialization limits, or answer contracts, a score gap should not automatically be attributed to better memory selection.
The evidence supports an audit method, not a universal memory policy
The study uses 24 source pairs, four for each of six mechanisms, inside one synthetic append-only event grammar. Futures contain only two update chunks. The selector directly understands that grammar. The generated baseline is not delivery-matched to deterministic selection, and the metamorphic repair analysis is post-hoc and reuses the original histories.
The offline transformation counts therefore should not be read as hundreds of independent benchmark examples, and the reported frontier-versus-structured identification intervals are fixed-sample bounds for unresolved outcomes, not population confidence intervals.
What the paper establishes more convincingly is methodological: present correctness and future sufficiency are different evaluation targets, and the gap can be made observable with controlled shared-future updates.
For persistent agents, that is enough to change the test. A memory system should not be certified only for the state it summarizes. It should be tested for the state transitions it is expected to survive.
Cognaptus: Automate the Present, Incubate the Future.
-
Guangzhe Zhang (2026). Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression. arXiv:2609.20045. https://arxiv.org/abs/2609.20045 ↩︎