跳到正文
arXiv:cs.LG· Ian de Holanda Cavalcanti Bezerra, Vivek Trivedy, Lucas Pascotti Valem, Longin Jan Latecki·· 4 小时前

面向细粒度图像检索的区域感知 CLS Token 增强方法

Region-Aware CLS Token Augmentation for Fine-Grained Image Retrieval

AI 导读

研究提出一种区域感知的语义 token 增强方法,在 DINOv2-reg 的全局 [CLS] token 和四个 register token 基础上匹配空间 patch 并提取 N x N 局部 ROI token,无需外部边界框或显著性模块即可自动定位感兴趣区域。

正文

View PDF HTML (experimental)

Abstract:Image retrieval methods often rely on a single global semantic descriptor extracted from an image, e.g., the [CLS] token in vision transformers. However, trying to squeeze all the semantic information of an image into a single descriptor can hurt downstream retrieval performance, especially for fine-grained retrieval tasks. In this work, we augment the semantic tokens in the newer visual transformers, the global [CLS] token and the four register tokens, with a carefully selected collection of spatial tokens, aiming to capture the spatial region representation that characterizes the contents captured in each of the semantic tokens. We leverage the DINOv2-reg model, which includes register tokens that emergently learn object and part-based representations. For each "cue" token ([CLS] and each register token), we find a "buddy" image patch token and extract an N x N patch region to produce a set of localized ROI tokens. Our approach automatically captures important regions of interest without any external bounding boxes or saliency modules, purely by matching semantic tokens with their spatial representation regions. Furthermore, we incorporate these tokens into a multi-vector retrieval framework inspired by ColBERT, enabling fine-grained matching via a per-token alignment mechanism while avoiding the large storage cost of keeping all patch embeddings. Through extensive experiments, we find that (1) register tokens encode useful fine-grained details that can complement the [CLS] token; (2) automatically pooled ROI tokens further improve fine-grained discrimination; and (3) multi-vector retrieval with a small set of tokens improves over a DINOv2-reg single-vector baseline while remaining tractable for large-scale search. The code is available at this https URL.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as: arXiv:2610.10991 [cs.CV]
  (or arXiv:2610.10991v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.10991

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Lucas Pascotti Valem [view email]
[v1] Wed, 7 Oct 2026 23:24:52 UTC (534 KB)

来源:arXiv:cs.LG · arxiv.org