跳到正文
arXiv:cs.LG· Vincent Siu, Glenn Grant-Richards, Vlad Pavlovich, Yizhou Sun, Dawn Song, Chenguang Wang·· 4 小时前AI 评分45

Transformer 拒答机制中的组件与维度稀疏性

Component and Dimension Sparsity in Transformer Refusal Mechanisms

AI 导读

研究将拒答激活引导分解为组件级干预,在四个开源权重模型中发现拒答方向集中于仅占上游组件 28–48% 的稀疏子集,仍保留 88–101% 的引导效果。在这些组件内,有效引导进一步集中在约 50% 的残差流维度,保留 85–98% 的效果,表明拒答由结构化、可识别的机制组装而成,而非弥散编码。代码与原始实验结果已公开,论文被 COLM 2026 接收。

正文

View PDF HTML (experimental)

Abstract:Activation steering manipulates large language model behavior by intervening on internal activations, but the mechanistic basis of these interventions remains poorly understood. We decompose refusal steering into component-level interventions across four open-weight models, identifying the sparse subsets of attention and MLP components whose steering suffices to reproduce the full behavioral effect. We find that refusal directions concentrate in sparse component mechanisms comprising 28--48\% of upstream components, retaining 88--101\% of steering effectiveness. Within these mechanisms, effective steering further concentrates in approximately 50\% of residual stream dimensions, retaining 85--98\% of the component-mechanism baseline, consistent with a privileged basis structure. Sparsity thus operates at two levels: which components are steered, and which dimensions within those components carry the signal. Together these findings show that refusal is not diffusely encoded across a transformer but assembled by a structured, identifiable mechanism, providing a foundation for mechanistic understanding of how refusal behaviors are represented and steered. To facilitate reproducibility, we release all code and raw experimental results in this https URL.
Comments: Accepted to COLM 2026
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.06903 [cs.CL]
  (or arXiv:2610.06903v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.06903

arXiv-issued DOI via DataCite

Submission history

From: Vincent Siu [view email]
[v1] Wed, 30 Sep 2026 23:32:40 UTC (5,599 KB)

来源:arXiv:cs.LG · arxiv.org