跳到正文
arXiv:cs.LG· Hubert Baniecki, Przemyslaw Biecek, Fabian Fumagalli·· 4 小时前AI 评分41

可解释性的元博弈与元归因:统一梯度与注意力方法的交互效应框架

The Metagame of Interpretability and Meta-Attributions

AI 导读

研究者提出"元博弈"(metagame)框架,将归因值 φᵢ 视作特征间合作博弈并计算其 Shapley 值,得到方向性元归因 φ_{j→i},可把任意梯度或注意力归因方法扩展到二阶交互效应。

正文

View PDF HTML (experimental)

Abstract:How can an arbitrary attribution method be generalized from first principles to capture interactions? We answer this with the metagame, a conceptual framework for quantifying second-order interaction effects of model explanations. We cast the attribution value $\phi_i$ of feature $i$ as a cooperative game among the other features and compute its Shapley value, which measures how much feature $j$ influences the attribution of $i$, yielding the directional meta-attribution $\varphi_{j \to i}$. By decomposing attribution itself rather than the model directly, meta-attributions extend any gradient- or attention-based method to interactions, uniting removal-based perturbations with model internals. Theoretically, we prove that meta-attributions sum to the first-order attribution they explain, a hierarchical decomposition that Shapley interactions and integrated Hessians turn out to perform implicitly. Empirically, we demonstrate that meta-attributions deliver insights across diverse interpretability applications: (i) quantifying token interactions in instruction-tuned language models, (ii) explaining cross-modal similarity in vision-language encoders, and (iii) interpreting text-to-image concepts in multimodal diffusion transformers.
Comments: NeurIPS 2026. Code: this https URL
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
Cite as: arXiv:2605.06295 [cs.LG]
  (or arXiv:2605.06295v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2605.06295

arXiv-issued DOI via DataCite

Submission history

From: Hubert Baniecki [view email]
[v1] Thu, 7 May 2026 13:59:26 UTC (4,411 KB)
[v2] Tue, 6 Oct 2026 16:56:18 UTC (4,112 KB)

来源:arXiv:cs.LG · arxiv.org