arXiv:cs.LG· Hadi Vafaii, Tejas Rao, David Chanin, Thomas Fel, Jacob L. Yates, Bruno Olshausen, David Klindt, Dileep George, Miguel L\'azaro-Gredilla·· 4 小时前AI 评分44
稀疏自编码器的推理与学习即自然梯度流:BeFOND 模型提出
Inference and learning in sparse autoencoders as natural gradient flow
AI 导读
BeFOND 是一种无编码器的稀疏编码模型,将推理与字典学习统一为共享变分自由能上的自然梯度流,具备闭式推理与学习动力学。在合成数据上,它提升了字典恢复与稀有特征检测,且随特征叠加程度增加对摊销基线的优势持续扩大;在语言模型激活上,它以更少训练数据在单特征概念检测与选择性干预上超过预训练参考 SAE,特征质量随字典宽度提升而持续改善,基线则基本趋于平台。
正文
Abstract:Sparse autoencoders are widely used to uncover interpretable features in neural networks, yet reliable recovery remains difficult when features overlap or activate infrequently. These challenges involve both inferring which features explain an input and learning the dictionary that represents them. Here, we unify inference and dictionary learning as natural-gradient flows on a shared variational free energy. We instantiate this framework as BeFOND, an encoder-free sparse coding model with closed-form inference and learning dynamics. We show how recurrent explaining away reduces interference between overlapping features, while Fisher preconditioning can compensate for the slow learning of rare features. On synthetic data, BeFOND improves dictionary recovery and rare-feature detection, with a growing advantage over amortized baselines as superposition increases. On language-model activations, it improves single-feature concept detection and selective intervention, outperforming pretrained reference SAEs with substantially less training data. Its feature quality continues to improve with dictionary width, whereas the evaluated baselines largely plateau. Together, these results show how improving inference and learning within a unified probabilistic framework can make better use of data and dictionary capacity to interpret and intervene on neural representations.
| Comments: | Code: this https URL |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07389 [cs.LG] |
| (or arXiv:2610.07389v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07389 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hadi Vafaii [view email]
[v1]
Mon, 5 Oct 2026 21:03:21 UTC (1,488 KB)
来源:arXiv:cs.LG · arxiv.org