arXiv:cs.LG(机器学习,全量分类)· Areeb Ahmad, Pratinav Seth, Vinay Kumar Sankarapu·· 14 小时前AI 评分41
语言模型自修复现象新解:消融即剂量,反权重机制研究
Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair
AI 导读
研究提出语言模型的"自修复"现象实为反权重(counterweight)在对比信号出现时的常规操作,其因果响应遵循仿射定律 E_r(λ)=own_r+γ_rλ。在 Gemma、Qwen、LLaMA、Mistral 四个模型的事实判定任务中,81 个下游方向有 68 个遵循该定律;GPT-2 Small 的 IOI 电路中 10 个注意力头有 7 个遵循,且全部为反权重。
正文
Abstract:Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear. The most systematic study to date concluded that self-repair is noisy and unlikely to have a single explanation. We argue that it has one: a gain already present before any ablation. Any intervention on a causally important component can be viewed as a point on a coordinate axis $\lambda$, the signed strength of a counterfactual contrast. Hence, conventional ablation methods are uncalibrated points on this axis. We show that the causal repair response for a fine-grained unit $r$ is governed by an affine law, $E_r(\lambda)=\mathrm{own}_r+\gamma_r\lambda$. The slope $\gamma_r$ is a fixed coefficient that consistently influences the model, with or without ablation, and its sign determines whether the unit counteracts or reinforces the removed signal. On a factual-verdict task across four models from distinct families (Gemma, Qwen, LLaMA, and Mistral), we identify components including MLP neurons, OV neurons, and singular directions that follow this affine law, 68 of 81 downstream directions in all. Moreover, we can anticipate the magnitude of $\gamma_r$ from the fixed weights. On the IOI circuit of GPT-2 Small, seven of the ten heads the intervention can reach follow the law, and all seven are counterweights. From this perspective, what may appear as self-repair is a counterweight performing its usual operation when the contrastive signal emerges at the core.
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.02173 [cs.LG] |
| (or arXiv:2610.02173v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02173 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Pratinav Seth [view email]
[v1]
Thu, 1 Oct 2026 17:57:04 UTC (1,001 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org