TL;DR for operators

A long recording may be queried many times after ingestion, while the model can expose only a limited number of frames directly in its prompt. The usual choices—sample harder, compress visual tokens, append memory tokens, or keep more model-internal cache—each spend capacity somewhere.

PReM changes where that historical evidence lives. It writes the video into a small recurrent state outside the prompt, then lets later questions retrieve from that state by steering the key and value representations of prompt positions already present. The frozen VLM still receives its normal bounded visual buffer; the recurrent memory supplements rather than replaces it.

In the paper’s strongest low-budget example, Qwen2.5-VL-3B rises from 50.39% to 53.45% macro-average accuracy at an offline visual budget of 16 frames. Gains remain positive across every tested visual budget in the principal matched-base comparisons. The standard four-slot memory state is about 256 KiB.

For searchable archives of meetings, procedures, training recordings, or operational footage, the practical proposition is write once and query many times. That proposition is promising but bounded: the experiments use 1-FPS ingestion, at most 240 writer frames, spatially pooled historical features, and primarily end-of-stream multiple-choice QA.

The scarce resource is not only context length

Suppose a two-hour training recording has already been processed. One user asks when a safety check occurred; another asks what happened immediately before a machine stopped; a third asks which step was repeated. Replaying or repacking the relevant history into a prompt for every question is expensive, yet retaining every frame directly is usually impossible under a fixed visual budget.

This is normally treated as a selection problem: decide which frames or compressed tokens deserve prompt space. The difficulty is that relevance depends on a question that may not exist when the video arrives. A brief event discarded during ingestion cannot be recovered later simply because a future query turns out to need it.

That is the design problem addressed by PReM, or Prefix-Steered Recurrent Memory.1 The paper’s most revealing result appears under a tight direct-visual budget. With Qwen2.5-VL-3B at offline $B=16$, the matched frozen base reaches 50.39% macro-average accuracy across six benchmarks; adding PReM raises that to 53.45%, a 3.06-point gain.

The result does not show that direct visual context is unnecessary. It shows that prompt allocation is not the only decision available.

Historical evidence stays outside the prompt

PReM separates video ingestion from question answering.

As frames arrive, the frozen visual encoder produces features that are spatially mean-pooled and projected into address and content representations. A query-agnostic writer updates a fixed number of external memory slots. Each slot is an associative matrix accompanied by a confidence value.

The write is residual rather than indiscriminate. For each incoming frame, the system estimates what a slot already predicts and writes mainly the remaining content. Routing, stability, novelty, visual salience, and confidence-protection gates control how strongly that residual modifies memory. In simplified form, the update is:

$$ S_m \leftarrow \sigma(f_m)S_m + \eta_w g_{t,m} e_{t,m}\kappa_{t,m}^{\top}. $$

The operational point is more important than the notation: a new frame should consume memory capacity primarily when it contributes information not already represented.

This state is not appended later as a bank of retrieved video tokens. The normal visual buffer remains in the prompt for direct grounding, while the recurrent state sits outside the sequence.

That distinction changes the cost structure. The standard four-slot recurrent state occupies about 256 KiB, and the paper reports roughly 0.03 GiB of additional peak GPU memory across visual budgets.

The question reads memory by steering existing prefix positions

Stored evidence still has to influence the frozen language model. PReM does this during prefill.

A representation of the non-visual prompt prefix—including the question—addresses the external memory. Retrieved values are routed across slots and projected into additive perturbations for the attention keys and values of existing non-visual prefix positions.

Conceptually:

$$ K_i^{(\ell)} = W_K^{(\ell)} \left(h_i^{(\ell)} + d_i^{k,(\ell)}\right), \qquad V_i^{(\ell)} = W_V^{(\ell)} \left(h_i^{(\ell)} + d_i^{v,(\ell)}\right). $$

The backbone projections remain frozen. No new memory tokens are appended, and the recurrent state is not queried repeatedly during autoregressive decoding. Once the steered prefix enters the ordinary KV cache, generation proceeds through the backbone’s normal decoding path.

This is the paper’s central architectural contribution: long-horizon evidence is stored externally, but its retrieval is expressed through representations the frozen model already knows how to use.

The diagnostics test whether the memory is doing real work

The headline benchmark gains are the main evidence. The ablations answer a different question: whether the architecture’s components explain those gains.

One diagnostic removes the decoder’s direct visual buffer entirely. With PReM memory alone, macro-average accuracy reaches 38.33%, compared with 34.82% for a vision-free base. That does not establish adequate video understanding without visible frames. It does provide direct evidence that the recurrent state retains task-relevant visual information rather than functioning only as a generic learned prompt adapter.

A second ablation separates key and value steering. At offline $B=16$ on Qwen2.5-VL-3B, the full design reaches 53.45%, versus 51.66% with key-only steering and 51.92% with value-only steering. The experiment supports the joint interface: retrieved memory appears to benefit from influencing both attention addressing and the content carried through attended values.

The write mechanism receives similar diagnostic support. Removing or simplifying individual gating terms lowers the same macro average to between 52.03% and 52.66%.

These are component tests, not additional independent benchmark theses. Their role is to make the performance result more interpretable.

Gains persist beyond one budget and one backbone

The matched comparisons are broader than the low-budget example.

At $B=90$, PReM raises offline macro-average accuracy from 56.06% to 57.95% on Qwen2.5-VL-3B and from 63.87% to 64.87% on Qwen3-VL-8B. In end-of-stream streaming evaluation, the corresponding gains are 1.09 and 1.43 percentage points.

The method also transfers to LLaVA-Video-7B. Three-seed experiments at offline $B=16$ show little training variation: Qwen2.5-VL-3B PReM averages 53.38% with a 0.06-point standard deviation, while Qwen3-VL-8B averages 59.11% with a 0.03-point standard deviation.

Duration diagnostics are consistent with the intended use case. On Video-MME, the reported gain increases from 0.67 points for short videos to 1.89 points for long videos. The longest LVBench group, containing videos of at least 60 minutes, shows a 5.61-point gain with a paired bootstrap confidence interval above zero.

The pattern is suggestive rather than universal: positive macro averages do not mean every benchmark entry improves under every configuration.

Write-once/query-many is the clearest deployment case

For an operator of a repeatedly queried video archive, the relevant system decision is where to pay for history.

Keeping more frames directly available spends prompt budget. Token-memory approaches also consume sequence positions when memory is retrieved. KV-cache approaches retain more model-internal activation state. PReM instead pays an ingestion cost to produce a compact reusable external state.

In the paper’s runtime experiment, ingestion adds roughly two seconds, while answer latency remains close to the frozen base. Cached recurrent states can be reused across questions; 41.4% of benchmark questions in the authors’ evaluation workflow shared cached states.

That creates a plausible business pathway when the same recording receives multiple later queries. The potential benefit is not simply higher benchmark accuracy. It is separating a one-time historical-ingestion cost from repeated answer-time use while keeping prompt-token consumption bounded.

The value proposition is weaker when every video is queried only once, when the required evidence is densely spatial, or when high-frequency actions dominate. Those workloads do not align as closely with what the experiments establish.

The evidence stops short of a continuous video agent

Several boundaries materially constrain interpretation.

Frames are ingested at 1 FPS, and streams longer than the evaluation horizon are uniformly subsampled to at most 240 writer frames. Historical memory is built from spatially mean-pooled frame features, so it does not preserve full intra-frame layout. Memory capacity is fixed rather than dynamically allocated.

Most importantly, the streaming experiments answer after the stream ends. They do not demonstrate arbitrary mid-stream questioning, continuous open-ended dialogue, or lifelong multimodal interaction.

The paper therefore supports a narrower conclusion: a small external recurrent state can improve frozen-VLM long-video QA under matched visual budgets without consuming additional memory-token positions. Whether the same interface can carry sufficiently detailed, adaptive memory for continuously operating video agents remains unresolved.

Memory placement becomes a system-design choice

PReM’s contribution is less about making the context window larger than about reconsidering where long-horizon information should live.

The experiments show that useful historical evidence can be compressed into a small external state, retrieved according to a later question, and injected into a frozen model without becoming another sequence of prompt tokens. The strongest practical fit is a recording that is expensive to revisit but likely to be queried repeatedly.

For those workloads, long-video architecture is no longer only a question of how many frames fit. It is also a question of which information should remain directly visible, which should persist externally, and how cheaply that external state can re-enter the model when the question finally arrives.

Cognaptus: Automate the Present, Incubate the Future.


  1. Siru Zhong and Qiongyan Wang and Xiaohui Lv and Yuzheng Zhuang and Shuai Tao and Wulong Liu and Haohuan Fu and Yuxuan Liang (2026). PREM: Prefix-Steered Recurrent Memory for Long-Video Understanding. arXiv:2609.23601. https://arxiv.org/abs/2609.23601 ↩︎