Cover image

Relevant Is Not Authorized: Put Identity Before Agent Memory Retrieval

TL;DR for operators Bio-MemArt1 addresses a problem that ordinary memory retrieval does not solve: a memory can be highly relevant to a query and still belong to the wrong user. Its intervention is deliberately narrow. Each persistent KV-memory block receives a biometric owner template. At query time, the current face or palmprint representation is compared with those templates. Memories that fail a calibrated similarity threshold are excluded before semantic retrieval begins. The original MemArt retrieval and generation machinery then operates only on the surviving memories. ...

October 1, 2026 · 7 min · Zelina
Cover image

Cache Hits Can Create Queues

TL;DR for operators A multi-worker LLM service can make a request cheaper by sending it to a worker that already holds much of the required prefix in KV cache. The problem is that repeatedly favoring that worker can create a queue. Spreading traffic more evenly reduces congestion but may throw away the computational advantage of prefix reuse. ...

September 5, 2026 · 7 min · Zelina
Cover image

Shrink the KV Cache, Miss the Bottleneck

TL;DR for operators You can cut KV-cache memory substantially and still leave end-to-end latency almost unchanged. That is the central practical message of Jiang et al.’s survey of serving-time KV-cache optimization.1 The literature does not point to one interchangeable family of “KV optimizations.” Different techniques intervene at different points in the serving system: some change when KV work executes, some change where KV state resides or moves, and some change how much state is represented or retained. ...

September 5, 2026 · 8 min · Zelina
Cover image

Choose the Bottleneck Before the KV Cache Strategy

TL;DR for operators Mamo, Kogiou, Yi, and Yu compare three representative approaches to managing the memory accumulated during LLM generation: keep the full cache on the GPU, permanently discard selected cached tokens, or retain the larger cache in CPU memory and fetch selected entries during decoding.1 Their benchmark finds no dominant strategy. ...

September 4, 2026 · 7 min · Zelina
Cover image

Not Every Layer Deserves the Same Cache

TL;DR for operators A compressed KV cache creates two separate decisions: which tokens to retain inside each Transformer layer, and how much of the total cache budget each layer should receive. The second decision is easy to hide behind uniform allocations or simple depth schedules, but the evidence here suggests that those rules can spend scarce GPU memory in the wrong places. ...

September 1, 2026 · 7 min · Zelina
Cover image

Full Stack, Not Full Panic: Why Agentic AI Needs Safety Above and KV Discipline Below

Full Stack, Not Full Panic: Why Agentic AI Needs Safety Above and KV Discipline Below Enterprise AI has entered its awkward teenage years. It wants to be autonomous, helpful, context-aware, cheap, safe, fast, auditable, and preferably not the reason the legal department starts drinking before lunch. That is a lot to ask from “just use a bigger model.” ...

June 9, 2026 · 15 min · Zelina
Cover image

Cache Me If You Can: Why LLM Benchmarks Need Contamination-Resistant Data

The benchmark score is not the product. The test pipeline is. Benchmarks used to feel like neutral scoreboards. A model sat down, answered questions, received a number, and everyone pretended the number meant generalization. That story became less charming once benchmark questions started appearing in the same public data oceans used to train the models being tested. ...

June 3, 2026 · 20 min · Zelina
Cover image

The KV Cache Is Not a Detail: Why LLM Compression Needs a Control Plane

Bandwidth is one of those infrastructure costs that looks boring until it becomes the product bottleneck. A retrieval-augmented assistant gets a long document. An agentic workflow accumulates tool traces. A support chatbot reuses a large system prompt and a customer-history prefix. The model may be fast enough, the GPUs may be expensive enough, and yet the user still waits. Not because the model is thinking harder. Because the system is moving state. ...

May 27, 2026 · 15 min · Zelina
Cover image

No Free Tokens: The New Economics of LLM Inference

Opening — Why this matters now For the last few years, AI strategy has been narrated as a model-quality story: bigger models, better benchmarks, longer context windows, more agents, more demos, more adjectives. That story was useful. It was also incomplete. The less glamorous reality is now arriving with the invoice attached. LLM systems are not merely models. They are production services that consume GPU memory, scheduling capacity, engineering attention, and operational patience. Once a business moves from a prototype to repeated daily use, the question changes from “Can the model answer?” to “Can the system answer reliably, cheaply, and repeatedly when real users arrive at inconvenient times?” ...

May 7, 2026 · 16 min · Zelina
Cover image

Packing Memory, Not Problems: How Short Clips Teach AI to Think Long in Video

Memory is usually the boring part of AI demos. The model gets the spotlight. The prompt gets the applause. The generated video either looks magical or embarrassingly haunted. Somewhere underneath, quietly paying the bill, sits the memory system. It decides what the model can still remember, what it must forget, and how much GPU memory gets sacrificed to the gods of temporal coherence. ...

March 28, 2026 · 20 min · Zelina