跳到正文
arXiv:cs.AI· Ashok Prasad Neupane, Dipan Bartaula, Ankit Belbase, Saugat Adhikari, Samip Ghimire, Saroj Poudel, Binod Bhattarai, Danda Pani Paudel·· 4 小时前AI 评分44

FORESIGHT:无需重训练即可为流式 VLM 规划未来感知

Foresight: planning future perception in streaming VLMs without retraining

AI 导读

FORESIGHT 是一种双流架构,用两个共享权重、输入编码器和 KV cache 的 Siamese LLM,让第二个 LLM 跑在流前面预判未来上下文并规划计算,从而在免训练条件下动态配置流式 VLM 的推理。

正文

View PDF HTML (experimental)

Abstract:Existing streaming vision-language models (VLMs) continuously perceive and reason over visual streams, but their computational pathways remain fixed throughout inference. Consequently, they cannot adapt computation to evolving scene dynamics, where different future events demand different levels and forms of perception. We show that streaming VLMs inherently possess the ability to anticipate the immediate future, and leverage this capability to dynamically configure future computation in a training-free manner. Realizing such anticipatory computation, however, is very challenging: future anticipation must be sufficiently reliable to guide computation, planning must run concurrently with streaming inference, and online reconfiguration must incur negligible overhead. To address these challenges, we introduce FORESIGHT, a dual-stream architecture comprising two Siamese LLMs with shared weights, input encoders, and KV cache. The first LLM continuously processes incoming tokens, while the second runs ahead of the stream to anticipate future context, plan future computation, and generate task responses without interrupting streaming inference. Each plan decides when to reason next, what to check then, and how densely to sample, keeping transient evidence separate from persistent control. The resulting computation plan is executed online through an efficient reconfiguration protocol with schemaguided decoding and lightweight diff-based updates, enabling dynamic adaptation with low overhead. With a frozen Qwen3-VL-8B backbone, FORESIGHT achieves 23.0 mean joint F1 on OmniPro Online evaluation beating strongest trained baseline by 9.5%, while improving the backbone by 6.7 on StreamingBench and 15.4 on OVO-Bench, with the largest gain of 18.7 when evidence arrives later in the video stream. Our source code will be made publicly available.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.03123 [cs.CV]
  (or arXiv:2610.03123v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.03123

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ashok Neupane [view email]
[v1] Fri, 2 Oct 2026 10:44:08 UTC (5,841 KB)

来源:arXiv:cs.AI · arxiv.org