Correct Now Is Not Enough: Testing Agent Memory for What Happens Next
TL;DR for operators An agent can have a compressed memory that answers the current question perfectly and is already inadequate for the next update. Guangzhe Zhang’s Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression1 demonstrates the problem with paired synthetic histories. On DeepSeek, the paper’s frontier selector and a latest-only variant both score 96/96 on current reads. After a later update makes older evidence relevant, frontier remains at 96/96, while latest-only falls to 32/96. ...