arXiv:cs.LG· Yao Xiao, Reuben Tan, Zhen Zhu, Yuqun Wu, Jianfeng Gao, Derek Hoiem·· 6 小时前AI 评分47
ReToken:用视觉检索 token 提升长上下文 VLM
ReToken: Improving Long-Context VLMs with Visual Retrieval Token
AI 导读
ReToken 是一种可学习的嵌入向量,能从 VLM 内部表示中提取检索信号,在预填充的 KV cache 中选出与查询相关的视觉 token,无需独立检索器或重新编码。
正文
Abstract:Long visual contexts challenge vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once can exceed GPU memory limits. We present RETOKEN, a single learnable embedding that extracts retrieval signals from the VLM's internal representations to select query-relevant visual tokens from the pre-filled KV cache. This enables retrieval within the answering VLM, without a separate retriever or re-encoding. Despite being trained on only a small image-QA dataset, RETOKEN generalizes across image and video benchmarks. On Visual Haystacks, it improves Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4 points (>20% relative). On LVBench, it transfers zero-shot to long video and improves Qwen3VL-8B by 8.0 points. Thanks to its lightweight design, both training and long-video inference fit on a single H100. Code is available at: this https URL
| Comments: | Accepted to NeurIPS 2026. Code: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.28627 [cs.CV] |
| (or arXiv:2607.28627v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.28627 arXiv-issued DOI via DataCite |
Submission history
From: Yao Xiao [view email]
[v1]
Thu, 30 Jul 2026 17:59:56 UTC (380 KB)
[v2]
Tue, 6 Oct 2026 22:51:13 UTC (387 KB)
来源:arXiv:cs.LG · arxiv.org