arXiv:cs.LG(机器学习,全量分类)· Chuqin Geng, Li Zhang, Haolin Ye, Mark Zhang, Luke Zhang, Xujie Si·· 14 小时前AI 评分41
机制可解释性存在目标层面的恢复缺口:干预式忠实度为何会选错电路
Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability
AI 导读
机制可解释性研究揭示了一种"目标层面恢复缺口":干预定义的忠实度指标可能偏好一个规模相同但行为复现更差的电路。在四个人类参考任务和 InterpBench 上,EAP、EAP-IG、ACDC 与 Edge-SP 的输出均出现该失败,重采样下 KL 在人类参考任务上对 9.4%-41.2% 的候选电路对产生错误排序。
正文
Abstract:Mechanistic interpretability aims to recover the internal computations responsible for model behavior. Progress in automated circuit discovery is often framed as a search problem: better attribution or optimization should identify better mechanisms. This assumes that the evaluation objective can recognize a better circuit once it is found. We show that intervention-defined faithfulness can instead prefer an equally sized circuit that reproduces the model's behavior less well, creating an objective-level recovery gap. Across four human-reference tasks and InterpBench, we compare validation faithfulness with behavior on held-out prompts under fixed ordinary resampling. The behavioral criterion is agreement with the intact model, including its mistakes, except on Greater-Than, where we use semantic accuracy. Controlled reference edits reveal misranking without any discovery algorithm, and outputs of EAP, EAP-IG, ACDC, and Edge-SP exhibit the same failure. Under resampling, KL misranks 9.4%-41.2% of candidate pairs across these methods on the human-reference tasks. We investigate context distortion as an explanation: replacing excluded signals changes the inputs on which retained components operate. Restoring selected signals from the recipient's intact-model execution repairs 96 of 100 persistent KL misrankings from the discovery pool on both validation and held-out prompts. The circuits and their original behavioral scores remain unchanged. These findings show why better discovery alone is insufficient when its objective rewards the wrong candidate.
| Comments: | 34 pages, 2 figures |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.02098 [cs.LG] |
| (or arXiv:2610.02098v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02098 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Chuqin Geng [view email]
[v1]
Thu, 1 Oct 2026 17:25:23 UTC (269 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org