跳到正文
arXiv:cs.AI· Youxing LI·· 6 小时前AI 评分39

PixelTriage:让多模态助手在打开图像前决定哪些检索记忆值得用像素

Decide Before You Look: Learning Which Retrieved Memories Deserve Pixels

AI 导读

PixelTriage 是一个置于检索之后的小模型插件,不生成文本,只读对话、简短笔记和每张检索图像的缩略图,预判其像素能带来多少增益。它由冻结的 27B 模型在合成记忆片段上标注训练。搭配 7B 回答模型时,在 M³Exam、DMV、MemEye 上仅用 11–23% 视觉 token 即无明显精度损失,DMV 上比打开全部图像快 2.9 倍。

正文

View PDF HTML (experimental)

Abstract:Multimodal assistants answer questions from long-term memories that contain images. After retrieval, each retrieved image reaches the answering model either as pixels, at about a thousand visual tokens per image, or as a stored text proxy that often misses the detail the question asks about. We find that the benefit of pixels usually comes from one or two retrieved memories, and that it can be predicted before the answering model runs, without reading any full-resolution image. In PixelTriage, a plug-in placed after retrieval, a small model that does not generate text reads the dialogue, a short note and a thumbnail of each retrieved memory and predicts how much its pixels would add. It is trained on synthetic memory episodes labeled by a frozen 27B model that answers each question with and without each memory's pixels. With a 7B answering model, PixelTriage lies on the accuracy--cost frontier of M$^3$Exam, DMV and MemEye and uses 11--23\% of the visual tokens without a significant loss of accuracy. On DMV it answers 2.9 times faster than opening all images. It outperforms retrieval order and uniform down-sizing at equal budgets and transfers to other memory systems and to a 397B answering model.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.07984 [cs.CV]
  (or arXiv:2610.07984v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.07984

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Youxing Li Rick [view email]
[v1] Tue, 6 Oct 2026 08:47:45 UTC (319 KB)

来源:arXiv:cs.AI · arxiv.org