跳到正文
arXiv:cs.LG· Adam Pardyl, Siddhartha Gairola, Sukrut Rao, Adam Wr\'obel, Bartosz Zieli\'nski, Bernt Schiele, Dawid Rymarczyk·· 4 小时前AI 评分34

DisParQ:面向可解释视觉基础模型的自监督部件概念

DisParQ: Self-Supervised Part Concepts for Interpretable Vision Foundation Models

AI 导读

DisParQ 能从冻结的纯视觉自监督骨干网络(如 DINOv2)中学习空间落地的离散概念表示,无需类别标签和语言监督。每个图像 patch 只分配到一个可学习原型概念,并用连续残差量化出离散属性,再由空间解码器重建骨干表示。在 ImageNet 线性探测上以 83.2% top-1 接近 DINOv2 教师模型,概念一致性优于语言对齐模型。

正文

View PDF HTML (experimental)

Abstract:Concept-based vision models represent images through an intermediate layer of human-inspectable concepts, so what a model relies on can be traced to those concepts. However, those models are often limited to fixed categories or depend on language to define their concepts. We introduce DisParQ (Discrete Parts with Quantized attributes), a method that learns spatially grounded, discrete concept representations from a powerful frozen vision-only self-supervised backbone. It requires no class labels and no language supervision. Each image patch is assigned to exactly one concept from a learnable prototype dictionary, and only a sparse subset of concepts may activate per image. To capture how each concept varies across images (e.g., the type of a "wheel"), we learn continuous residuals alongside the concepts and then quantize them into discrete attributes. A spatial decoder reconstructs the backbone's representation from the concepts and attributes alone, so successful reconstruction means that the discrete representation preserves the backbone's information. We evaluate DisParQ across seven datasets, from general recognition (ImageNet, PartImageNet, Places) to fine-grained benchmarks (CUB, Cars, Dogs, Flowers). We show that DisParQ closely matches its frozen DINOv2 teacher on ImageNet linear probing (83.2% top-1), achieves higher concept consistency than language-aligned models, remains competitive on fine-grained recognition, and enables cross-category part-based retrieval.
Comments: Under review
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.09802 [cs.CV]
  (or arXiv:2610.09802v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.09802

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Dawid Rymarczyk Dr [view email]
[v1] Wed, 7 Oct 2026 10:19:04 UTC (16,701 KB)

来源:arXiv:cs.LG · arxiv.org