Cover image

Don’t Run the Whole Model Yet: ECHO Reallocates Verification Across Transformer Depth

TL;DR for operators Speculative decoding saves time when several candidate tokens can be checked together, but the verification step can become its own cost center: if every speculative cycle still traverses the full model, acceleration is bounded by how often that expensive check must run. ECHO1 changes where that verification work happens. It lets cheaper early Transformer layers screen and extend candidates repeatedly, then invokes the remaining layers less frequently for authoritative verification. Intermediate states are reused rather than recomputed. In the paper’s primary comparison table, ECHO reports overall speedups from 2.42× to 2.90× across five model configurations, with overall mean accepted tokens (MAT) from 4.16 to 5.55. ...

October 4, 2026 · 7 min · Zelina
Cover image

The Guard Is Already in the Draft: Reusing Speculative Decoding for LLM Monitoring

TL;DR for operators Production monitoring often creates an unattractive choice. A small probe is cheap enough to run on every request but may compress away evidence that occurred earlier in a sequence. A stronger position-aware classifier preserves more information but adds computation to an inference path that is already expensive. Speculative Probing: LLM Monitoring at Speculative-Decoding Cost1 asks whether part of that cost has already been paid. Some LLM deployments use a smaller auxiliary prediction component to accelerate generation. Speculative Probing reuses that component—its speculative-decoding head—as a frozen feature extractor for monitoring. The base model and draft head remain unchanged; only a few learned task-specific query vectors and a small classifier are trained. ...

September 28, 2026 · 8 min · Zelina
Cover image

Turn Down the Work, Not Just the Clock: Governing On-Device LLM Power Across Hardware and Decoding

TL;DR for operators Lowering an edge processor’s clock can reduce power, but it also slows token generation. That leaves a local LLM runtime with a constrained control problem: keep responses fast enough, stay below thermal limits, limit energy use, and avoid degrading output quality more than the application can tolerate. PELM1 expands the control surface instead of optimizing processor frequency alone. Its runtime governor adjusts CPU and GPU frequencies together with how the model drafts and verifies tokens. On the tested LLaMA-13B workload on Jetson AGX Orin, the paper reports 29.0% to 52.4% lower energy than the evaluated comparison methods across cooling conditions, while PELM delivered the highest average generation speed in three of four conditions. ...

September 28, 2026 · 7 min · Zelina
Cover image

Grow While You Roll: Training the Draft Head Inside the RL Run

TL;DR for operators LLM reinforcement learning spends a large share of its wall-clock budget generating responses. Speculative decoding can reduce that cost, but it normally assumes that a smaller predictor is already capable of proposing useful tokens before the main model verifies them. GrowMTP1 tests whether that capability can instead be learned during the RL workload itself. On Qwen3-4B, a randomly initialized five-token draft head reached an acceptance length of 2.91 after 500 mathematical-reasoning RL steps. Rollout generation became 2.13x faster, and the complete RL step—including draft-head training—became 1.60x faster. ...

September 26, 2026 · 7 min · Zelina
Cover image

Sparse Per Token, Dense Per Batch: XShare Reprices MoE Routing at Inference Time

TL;DR for operators Sparse MoE routing does not guarantee a sparse serving batch. As more tokens are decoded together, especially under speculative decoding, their individual top-k choices can collectively touch a large fraction of the model’s experts. That increases expert-weight movement during a memory-bound stage of inference. XShare treats this as a serving-policy problem rather than a model-retraining problem. It ranks experts using router scores from the current batch, limits the shared expert pool, and then lets each token perform its normal top-k refinement within that pool. In the reported GPT-OSS 120B experiments, moderate restrictions improve output-token throughput by roughly 6-13% in standard decoding and around 13-14% in several speculative settings while keeping accuracy comparatively close to baseline. More aggressive restriction can go much faster, but the quality loss can be large. ...

September 12, 2026 · 7 min · Zelina
Cover image

State of Delay: KVBuffer and the Memory Tax of Linear Attention

Latency has a habit of hiding inside words that sound efficient. “Constant decoding cost” is one of those phrases. It suggests a clean engineering promise: linear attention avoids the context-length explosion of softmax attention, so long-context inference should become simpler, cheaper, and less melodramatic. Very nice. The GPU accountants, however, have not retired. ...

June 6, 2026 · 15 min · Zelina
Cover image

The Tail That Wags the Model: Why p99 Latency Should Run Your LLM

A demo can survive a slow answer. A production service cannot survive the slow answer that arrives just often enough to make users stop trusting the product. That is the quiet problem behind p99 latency. The average response time tells you how the service feels on a normal day. p99 tells you what happens to the unlucky one percent: the support agent waiting in front of a customer, the analyst refreshing a dashboard, the employee whose workflow now includes watching a spinner and reconsidering their life choices. ...

March 15, 2026 · 14 min · Zelina
Cover image

Speculation, But With Standards: Training Draft Models That Actually Get Accepted

Queue. That is still the least glamorous word in AI infrastructure, and probably the most honest one. A user asks a model to write code, summarize a filing, inspect an image, or reason through a customer ticket. The model knows what to do, more or less. The bottleneck is not ambition. It is waiting: one token after another, one expensive forward pass after another, while the GPU performs a very sophisticated version of typing slowly. ...

February 8, 2026 · 13 min · Zelina
Cover image

Speculate Smarter, Not Harder: Hierarchical Decoding Without Regret

Speed is the polite word. Cost is the less polite one. Every production LLM system eventually meets the same boring villain: the target model must generate tokens one after another, and each forward pass is expensive. Speculative decoding was supposed to soften that problem. Let a cheaper draft model run ahead, ask the expensive model to verify the draft, and accept several tokens per target-model call when the draft is good enough. Simple. Elegant. Almost suspiciously useful. ...

January 12, 2026 · 16 min · Zelina