Cover image

Cache Hits Can Create Queues

TL;DR for operators A multi-worker LLM service can make a request cheaper by sending it to a worker that already holds much of the required prefix in KV cache. The problem is that repeatedly favoring that worker can create a queue. Spreading traffic more evenly reduces congestion but may throw away the computational advantage of prefix reuse. ...

September 5, 2026 · 7 min · Zelina
Cover image

Kernel Kombat: How Multi‑Agent LLMs Squeeze 1.32× More From Your GPUs

Kernel Kombat: How Multi-Agent LLMs Squeeze 1.32× More From Your GPUs GPU bills have a charming way of turning “just one more model deployment” into a finance meeting. For companies running large language model serving stacks, the problem is rarely that nobody knows GPUs matter. Everyone knows. The harder problem is that performance bottlenecks often live inside kernels most executives will never see: attention merges, normalization fusions, activation multiplications, tiny pieces of code called millions or billions of times until “small inefficiency” becomes “why is the infrastructure budget wearing a crown?” ...

September 13, 2025 · 14 min · Zelina