跳到正文
arXiv:cs.LG· Yusong Zhao, Hengyi Wang, Tanuja Ganu, Akshay Nambi, Hao Wang·· 4 小时前AI 评分39

SAE++:级联稀疏自编码器学习多模态 LLM 的多层级视觉概念

SAE++: Cascaded Sparse Autoencoders Learn Multi-Level Visual Concepts in Multimodal LLMs

AI 导读

SAE++ 是一种级联稀疏自编码器架构,通过在初级 SAE 的解码器权重上训练二级 SAE,直接学习"概念的概念",从而为多模态 LLM 构建层级化视觉概念。该方法避免了嵌套式共享前缀耦合与简单堆叠 SAE 的瓶颈,在 Qwen3-VL、Gemma-3 和 LLaVA 上的实验显示,其层级概念一致性优于现有 SAE 基线。

正文

View PDF HTML (experimental)

Abstract:Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language tasks, yet their internal visual representations remain difficult to interpret. Sparse Autoencoders (SAEs) provide a scalable way to decompose dense model activations into sparse, interpretable features. However, existing SAE architectures primarily recover flat feature dictionaries and are less suited for explicit multi-level concept organization. In this paper, we introduce a cascaded sparse autoencoder architecture, dubbed SAE++, for learning hierarchical visual concepts in MLLMs. Rather than nesting or stacking SAE sparse activation codes, SAE++ trains a second-level SAE directly on the decoder weights of the first-level SAE, treating learned low-level feature directions as inputs for higher-level abstraction. This design enables SAE++ to learn "concepts of concepts" while avoiding drawbacks from the shared-prefix coupling of nesting, Matryoshka-style hierarchies and the bottlenecks of naively stacked SAEs. Experiments across Qwen3-VL, Gemma-3, and LLaVA on multiple visual datasets show that SAE++ improves interpretability in terms of hierarchical concept coherence over state-of-the-art SAE baselines. Results on concept steering further demonstrate that the learned concept groups support effective group-level interventions in MLLM outputs. Code is available at this https URL.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2606.16193 [cs.CV]
  (or arXiv:2606.16193v2 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2606.16193

arXiv-issued DOI via DataCite

Submission history

From: Yusong Zhao [view email]
[v1] Mon, 15 Jun 2026 04:10:40 UTC (5,005 KB)
[v2] Tue, 6 Oct 2026 04:40:58 UTC (3,180 KB)

来源:arXiv:cs.LG · arxiv.org