跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Juli Huang·· 14 小时前AI 评分40

GPT-2 注意力头消融何时能支撑因果结论?投影层混淆、地板效应与匹配对照

When Do Attention-Head Ablations Support Causal Claims? Projection-Level Confounds, Floor Effects, and Matched Controls

AI 导读

研究用 GPT-2 small 证明,注意力头消融的因果推断依赖干预语义、评估指标与对照设计。投影后置的"置零"实现与修正的投影前置消融几乎不相关(Pearson r = 0.057),选出的 top-5 重要头完全不同;二值准确率会掩盖地板与天花板效应,而 gold-token 对数概率仍保持梯度。

正文

View PDF HTML (experimental)

Abstract:Attention-head ablation, zeroing a head and measuring the resulting change in task performance, is a common method for inferring which components of a language model are causally responsible for a behavior. We show using GPT-2 small that this inference can be fragile unless the intervention semantics, evaluation metric, and controls are carefully validated. A natural post-projection implementation of "zeroing a head" is nearly uncorrelated with a corrected pre-projection ablation (Pearson r = 0.057) and selects a completely disjoint top-5 set of important heads. We also show that binary accuracy can hide effects at behavioral floors and near ceilings, whereas gold-token log-probability remains graded. Using a discovery/held-out split and 1,000 matched random-head and layer-matched-head control draws, the corrected per-head effect ranking is highly stable across splits (Spearman rho = 0.974), and the top-5 selected heads significantly exceed both control distributions (Monte Carlo p = 0.001). However, evidence for task specificity is not robust on GPT-2. Replication on DistilGPT2 preserves the intervention-semantic and matched-control findings. These results show that single-head ablation does not by itself justify a causal claim; defensible interpretation requires correct intervention placement, a non-saturated continuous metric, and matched held-out controls.
Comments: Code available in the accompanying repository. 2 figures
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as: arXiv:2610.00373 [cs.LG]
  (or arXiv:2610.00373v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00373

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Julia Huang [view email]
[v1] Wed, 30 Sep 2026 07:28:42 UTC (140 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org