Cover image

Not Every Layer Deserves the Same Cache

TL;DR for operators A compressed KV cache creates two separate decisions: which tokens to retain inside each Transformer layer, and how much of the total cache budget each layer should receive. The second decision is easy to hide behind uniform allocations or simple depth schedules, but the evidence here suggests that those rules can spend scarce GPU memory in the wrong places. ...

September 1, 2026 · 7 min · Zelina
Cover image

State of Delay: KVBuffer and the Memory Tax of Linear Attention

Latency has a habit of hiding inside words that sound efficient. “Constant decoding cost” is one of those phrases. It suggests a clean engineering promise: linear attention avoids the context-length explosion of softmax attention, so long-context inference should become simpler, cheaper, and less melodramatic. Very nice. The GPU accountants, however, have not retired. ...

June 6, 2026 · 15 min · Zelina