Cover image

Long Context Without a Long Chain

TL;DR for operators Serving longer contexts usually creates an uncomfortable architectural choice. Attention keeps training and inference highly parallel but makes pairwise token interaction increasingly expensive as the sequence grows. Conventional recurrence reduces that interaction cost but makes later positions wait through a sequential chain. The paper studies a third computation pattern. Its AR-GRC model organizes recurrence as a balanced tree, producing all autoregressive prefix representations with $O(\log N)$ computational depth and $O(N)$ total composition work. In the reported measurements, evaluation time grows approximately linearly with context length. The model also preserves stable perplexity when moved from a maximum training context of 512 tokens to evaluation sequences as long as 2,976 tokens. ...

October 10, 2026 · 7 min · Zelina