arXiv:cs.AI· Roy Xie, Dan Friedman, Donghan Yu, Bowen Pan, Christopher Fifty, Jang-Hyun Kim, Xianzhi Du, Zhe Gan, Vivek Rathod, Bhuwan Dhingra·· 4 小时前AI 评分37
LensVLM:面向文本压缩视觉表征的选择性上下文扩展
LensVLM: Selective Context Expansion for Compressed Visual Representation of Text
AI 导读
基于 Qwen3.5-9B-Base 的 LensVLM 通过让 VLM 扫描压缩图像并按需用学习到的工具选择性展开相关图像,在 4.3 倍有效压缩下保持与全文上限相当的准确率,并在七个文本 QA 基准上以最高 10.1 倍有效压缩超越基于检索、文本压缩和视觉压缩的基线。
正文
Abstract:Vision Language Models (VLMs) offer the exciting possibility of processing text as rendered images, bypassing the need for tokenizing the text into long token sequences. Since VLM image encoders map fixed-size images to a fixed number of visual tokens, varying rendering resolution provides a fine-grained compression knob. However, accuracy deteriorates quickly as compression increases: characters shrink below the vision encoder's effective resolution, making them indistinguishable. To address this, we propose LensVLM, an inference framework and post-training recipe that enables VLMs to scan compressed images, then selectively expand only the relevant images to their uncompressed form via learned tools. Building on Qwen3.5-9B-Base, LensVLM maintains accuracy comparable to the full-text upper bound at 4.3$\times$ effective compression and outperforms retrieval-based, text- and visual-compression baselines up to 10.1$\times$ effective compression across seven text QA benchmarks. LensVLM also generalizes to multimodal document and code understanding tasks, with the accuracy gain over baselines growing as compression increases. Our analysis validates this approach: training makes visual compression robust to rendering choices, and as compression grows the model increasingly relies on expanded content rather than unreliable visual reading. The analysis also yields practical tool-choice guidance: text expansion is preferable for rendered text, while high-resolution image expansion suits native documents whose layout cues carry task-relevant information.
| Comments: | Accepted to NeurIPS 2026 |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2605.07019 [cs.CV] |
| (or arXiv:2605.07019v3 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2605.07019 arXiv-issued DOI via DataCite |
Submission history
From: Roy Xie [view email]
[v1]
Thu, 7 May 2026 23:03:21 UTC (14,384 KB)
[v2]
Thu, 1 Oct 2026 04:51:40 UTC (14,390 KB)
[v3]
Fri, 2 Oct 2026 04:38:07 UTC (14,390 KB)
来源:arXiv:cs.AI · arxiv.org