arXiv:cs.LG· Nils Grandien, David Steinmann, Felix Friedrich, Kristian Kersting·· 7 小时前AI 评分41
稀疏自编码器能否学到有意义的概念层级?
Do Sparse Autoencoders Learn Meaningful Concept Hierarchies?
AI 导读
一项研究为无监督概念发现中的泛化/特化层级推导出关键要求与具体评测协议,并将其应用于视觉数据上训练的现有 SAE 方法。结果显示,特征空间通常具备形成合理层级的基础,但建立良好的层级结构仍然困难,特征吸收(硬性与连续的软性形式)会系统性损害层级质量。
正文
Abstract:Sparse autoencoders (SAEs) have become an important tool for unsupervised concept discovery in large models. To make the resulting feature spaces more interpretable and manageable, recent approaches have begun imposing hierarchical structure, either explicitly or as an implicit effect of training constraints, yet rigorous comparison remains difficult. There are no agreed-upon requirements for what a meaningful feature hierarchy should satisfy, and evaluation has largely relied on qualitative illustrations with fragmented quantitative protocols. To address this, we derive a set of key requirements for generalization/specialization hierarchies in unsupervised concept discovery, drawing on semantic net and taxonomy research alongside recent SAE work, and use them to derive a concrete evaluation protocol. Applying this protocol to current SAE approaches trained on visual data, we find that while feature spaces generally provide a basis for sensible hierarchies, establishing good hierarchical structure remains challenging. In particular, feature absorption, both in its well-known hard form and in a continuous, soft form, systematically compromises hierarchy quality, pointing to a fundamental tension that future approaches will need to navigate.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2606.22994 [cs.LG] |
| (or arXiv:2606.22994v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2606.22994 arXiv-issued DOI via DataCite |
Submission history
From: David Steinmann [view email]
[v1]
Mon, 22 Jun 2026 08:12:34 UTC (10,451 KB)
[v2]
Tue, 6 Oct 2026 15:34:38 UTC (10,514 KB)
来源:arXiv:cs.LG · arxiv.org