跳到正文
arXiv:cs.AI· Sanket Gandhi, Utkarsh Giri, Varun Subramanium, Rohan Paul, Parag Singla·· 3 小时前

LinSlot:利用线性表示假说从槽位对象表示中无监督发现属性

LinSlot: Exploiting Linear Representation hypothesis for unsupervised attribute discovery from slot based object representation

AI 导读

研究者提出 LinSlot,一种联合发现对象与属性表示的概率框架,利用线性表示假说(LRH)将可组合概念建模为槽位表示中的线性可加子空间。该架构通过 block attention 将属性表示连接到槽位,并在对象和属性两个表示空间中引入 LRH,从而优化所提图模型的 ELBO。在多个数据集上,LinSlot 的 DCI 分数优于当前最优方法,并支持解耦可解释表示带来的图像编辑能力。

正文

View PDF HTML (experimental)

Abstract:This paper studies the problem of learning disentangled representations of objects and their attributes from raw, unstructured image data. Slot-based methods have shown considerable success in unsupervised learning of object representations from images. Block-slot attention-based methods extend this framework to attribute representations by assuming a uniform factorization of object representations into attributes, which may be suboptimal and consequently limit the quality of the learned representations. We therefore investigate a framework for jointly discovering object and attribute representations. Our key contribution is leveraging the Linear Representation Hypothesis (LRH), which postulates that composable concepts can be represented as linearly additive subspaces in slot representations. Based on this insight, we propose a probabilistic model connecting images, slots (objects), and blocks (attributes). We present an architecture that leverages block attention to connect attribute representations to slots and incorporates LRH in both object and attribute representation spaces. This architecture effectively optimizes the Evidence Lower Bound (ELBO) of the proposed graphical model. Our experiments demonstrate (i) effective discovery of disentangled object and attribute representations, (ii) empirical evidence for LRH in slot space, and (iii) the ability to perform image editing owing to the disentangled and interpretable nature of the learned representations. Our experiments on multiple datasets demonstrate improvements in DCI scores over state-of-the-art methods.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.10722 [cs.CV]
  (or arXiv:2610.10722v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.10722

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sanket Gandhi [view email]
[v1] Wed, 7 Oct 2026 18:04:19 UTC (1,215 KB)

来源:arXiv:cs.AI · arxiv.org