Cover image

Don’t Run the Whole Model Yet: ECHO Reallocates Verification Across Transformer Depth

TL;DR for operators Speculative decoding saves time when several candidate tokens can be checked together, but the verification step can become its own cost center: if every speculative cycle still traverses the full model, acceleration is bounded by how often that expensive check must run. ECHO1 changes where that verification work happens. It lets cheaper early Transformer layers screen and extend candidates repeatedly, then invokes the remaining layers less frequently for authoritative verification. Intermediate states are reused rather than recomputed. In the paper’s primary comparison table, ECHO reports overall speedups from 2.42× to 2.90× across five model configurations, with overall mean accepted tokens (MAT) from 4.16 to 5.55. ...

October 4, 2026 · 7 min · Zelina