arXiv:cs.LG· Tue M. Cao, Lisiane Pruinelli, My T. Thai·· 5 小时前AI 评分40
ASPD:基于激活空间的权重读写特征可扩展参数分解
Weights Read and Write Features: Scalable Parameter Decomposition Grounded in Activation Space
AI 导读
研究者提出 Activation-Supported Parameter Decomposition(ASPD),联合分解激活空间与参数空间,将每个学到的权重组件锚定到其读取或写入的激活特征上,并在 Qwen-3-8B 上验证。
正文
Abstract:Activation space and parameter space provide complementary views of model computation. Activations represent information, while weights read, transform, and write that information. Yet existing interpretability methods largely study the two spaces separately, leaving the connection between represented information and parameter-level computation underexplored. We introduce Activation-Supported Parameter Decomposition (ASPD), which jointly decomposes activation and parameter spaces and grounds each learned weight component in the activation features it reads or writes. This grounding constrains otherwise non-unique parameter decompositions using the model's internal activations, while an internal reconstruction objective provides a local learning signal at the weight matrix being analyzed. Together, these properties enable scalable, interpretable, and causally editable parameter decomposition in pretrained large language models, demonstrated on Qwen-3-8B. The learned read--write components can also be composed into parameter-level mechanism circuits. We use ASPD to recover mechanisms underlying the classic IOI circuit and trace semantic transformations through model weights.
| Comments: | preprint |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.37731 [cs.LG] |
| (or arXiv:2609.37731v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.37731 arXiv-issued DOI via DataCite |
Submission history
From: Tue Minh Cao [view email]
[v1]
Tue, 29 Sep 2026 14:53:32 UTC (16,947 KB)
[v2]
Thu, 1 Oct 2026 19:33:39 UTC (16,947 KB)
来源:arXiv:cs.LG · arxiv.org