跳到正文
arXiv:cs.LG· Yao Xiao, Reuben Tan, Zhen Zhu, Yuqun Wu, Jianfeng Gao, Derek Hoiem·· 6 小时前AI 评分47

ReToken:用视觉检索 token 提升长上下文 VLM

ReToken: Improving Long-Context VLMs with Visual Retrieval Token

AI 导读

ReToken 是一种可学习的嵌入向量,能从 VLM 内部表示中提取检索信号,在预填充的 KV cache 中选出与查询相关的视觉 token,无需独立检索器或重新编码。

正文

View PDF HTML (experimental)

Abstract:Long visual contexts challenge vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once can exceed GPU memory limits. We present RETOKEN, a single learnable embedding that extracts retrieval signals from the VLM's internal representations to select query-relevant visual tokens from the pre-filled KV cache. This enables retrieval within the answering VLM, without a separate retriever or re-encoding. Despite being trained on only a small image-QA dataset, RETOKEN generalizes across image and video benchmarks. On Visual Haystacks, it improves Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4 points (>20% relative). On LVBench, it transfers zero-shot to long video and improves Qwen3VL-8B by 8.0 points. Thanks to its lightweight design, both training and long-video inference fit on a single H100. Code is available at: this https URL
Comments: Accepted to NeurIPS 2026. Code: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2607.28627 [cs.CV]
  (or arXiv:2607.28627v2 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2607.28627

arXiv-issued DOI via DataCite

Submission history

From: Yao Xiao [view email]
[v1] Thu, 30 Jul 2026 17:59:56 UTC (380 KB)
[v2] Tue, 6 Oct 2026 22:51:13 UTC (387 KB)

来源:arXiv:cs.LG · arxiv.org