Cover image

Store the Video Once, Steer the Prompt Later

TL;DR for operators A long recording may be queried many times after ingestion, while the model can expose only a limited number of frames directly in its prompt. The usual choices—sample harder, compress visual tokens, append memory tokens, or keep more model-internal cache—each spend capacity somewhere. PReM changes where that historical evidence lives. It writes the video into a small recurrent state outside the prompt, then lets later questions retrieve from that state by steering the key and value representations of prompt positions already present. The frozen VLM still receives its normal bounded visual buffer; the recurrent memory supplements rather than replaces it. ...

October 11, 2026 · 8 min · Zelina