arXiv:cs.LG· Hubert Baniecki, Przemyslaw Biecek, Fabian Fumagalli·· 4 小时前AI 评分41
可解释性的元博弈与元归因:统一梯度与注意力方法的交互效应框架
The Metagame of Interpretability and Meta-Attributions
AI 导读
研究者提出"元博弈"(metagame)框架,将归因值 φᵢ 视作特征间合作博弈并计算其 Shapley 值,得到方向性元归因 φ_{j→i},可把任意梯度或注意力归因方法扩展到二阶交互效应。
正文
Abstract:How can an arbitrary attribution method be generalized from first principles to capture interactions? We answer this with the metagame, a conceptual framework for quantifying second-order interaction effects of model explanations. We cast the attribution value $\phi_i$ of feature $i$ as a cooperative game among the other features and compute its Shapley value, which measures how much feature $j$ influences the attribution of $i$, yielding the directional meta-attribution $\varphi_{j \to i}$. By decomposing attribution itself rather than the model directly, meta-attributions extend any gradient- or attention-based method to interactions, uniting removal-based perturbations with model internals. Theoretically, we prove that meta-attributions sum to the first-order attribution they explain, a hierarchical decomposition that Shapley interactions and integrated Hessians turn out to perform implicitly. Empirically, we demonstrate that meta-attributions deliver insights across diverse interpretability applications: (i) quantifying token interactions in instruction-tuned language models, (ii) explaining cross-modal similarity in vision-language encoders, and (iii) interpreting text-to-image concepts in multimodal diffusion transformers.
| Comments: | NeurIPS 2026. Code: this https URL |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML) |
| Cite as: | arXiv:2605.06295 [cs.LG] |
| (or arXiv:2605.06295v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.06295 arXiv-issued DOI via DataCite |
Submission history
From: Hubert Baniecki [view email]
[v1]
Thu, 7 May 2026 13:59:26 UTC (4,411 KB)
[v2]
Tue, 6 Oct 2026 16:56:18 UTC (4,112 KB)
来源:arXiv:cs.LG · arxiv.org