arXiv:cs.LG· Vit\'oria Barin-Pacela, Shruti Joshi, Isabela Camacho, Simon Lacoste-Julien, David Klindt·· 3 小时前
为何线性探针与稀疏自编码器在组合泛化上失败
Stop Probing, Start Coding: Why Linear Probes and Sparse Autoencoders Fail at Compositional Generalisation
AI 导读
SAE 的摊销间隙在训练集规模、隐维度与稀疏度各档下持续存在,导致其在分布外组合迁移中失效。研究将失败归因于字典学习而非推理过程:SAE 学到的字典方向明显错误,用逐样本 FISTA 替换编码器也无法弥合差距。在 Pythia-70M 和 Gemma-2-2B 的 LLM 激活上,oracle 基线证明只要字典足够好,问题在所有测试规模下均可解。
正文
Abstract:The linear representation hypothesis states that neural network activations encode high-level concepts as linear mixtures. However, under superposition, this encoding is a projection from a higher-dimensional concept space into a lower-dimensional activation space, and a linear decision boundary in the concept space need not remain linear after projection. In this setting, classical sparse coding methods with per-sample iterative inference leverage compressed sensing guarantees to recover latent factors. Sparse autoencoders (SAEs), on the other hand, amortise sparse inference into a fixed encoder, introducing a systematic gap. We show this amortisation gap persists across training set sizes, latent dimensions, and sparsity levels, causing SAEs to fail under out-of-distribution (OOD) compositional shifts. Through controlled experiments that decompose the failure, we identify dictionary learning as the limiting factor (not the inference procedure): SAE-learned dictionaries point in substantially wrong directions, and replacing the encoder with per-sample FISTA on the same dictionary does not close the gap. An oracle baseline proves the problem is solvable with a good dictionary at all scales tested. Our results, including experiments with real LLM activations (Pythia-70M, Gemma-2-2B) reframe the SAE failure as a dictionary learning challenge, not an inference problem, and point to scalable dictionary learning as the key open problem for sparse inference under superposition.
| Comments: | Accepted for publication at UAI 2026. This is a slightly updated version of the published manuscript; see Corrigendum at the end of the paper |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2603.28744 [cs.LG] |
| (or arXiv:2603.28744v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2603.28744 arXiv-issued DOI via DataCite |
|
| Journal reference: | Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:364-412, 2026 |
Submission history
From: Vitoria Barin-Pacela [view email]
[v1]
Mon, 30 Mar 2026 17:52:16 UTC (2,880 KB)
[v2]
Wed, 7 Oct 2026 17:25:17 UTC (4,315 KB)
来源:arXiv:cs.LG · arxiv.org