arXiv:cs.LG· Thomas Schnake, Doreen Sch\"oppenthau, Alexander Meyer, Jacques Corbeil, Klaus-Robert M\"uller, Gr\'egoire Montavon·· 4 小时前AI 评分35
通过模型无关的概念字典重新审视可解释 AI
Revisiting Explainable AI through Model-Independent Concept Dictionaries
AI 导读
研究者提出 DictXAI,一种在输入域中通过字典直接定义概念的可解释 AI 方法,先计算输入的稀疏编码,再将模型预测归因于相应字典元素。该方法能将 AI 故障(如 Clever Hans 效应)直接归因于数据中可识别的伪影模式,并可跨学习图像基、心电图解析波形及实验获取的字典元素等多种字典运作。
正文
Abstract:Modern applications of AI rely on increasingly complex models. Explainable AI (XAI) has emerged as a set of techniques aimed at improving model transparency. However, existing XAI methods typically assume input features to be inherently interpretable, or they rely on intermediate internal abstractions that are difficult to characterize and highly architecture-specific, hindering consistent use across models. To address these limitations, we propose DictXAI, a method that defines concepts directly in the input domain via a dictionary---a large, potentially overcomplete set of predefined elements, each carrying an interpretable meaning. Technically, DictXAI first computes a sparse code of the input and then attributes the model's prediction to the associated dictionary elements. We demonstrate the actionable nature of DictXAI explanations, showing that they can attribute AI malfunctions (e.g., Clever Hans effects) directly to identifiable artifact patterns in the data, while fostering human-AI alignment on intricate biomedical signals. We further demonstrate our method's ability to operate across a wide variety of dictionaries, including learned image bases, analytically defined waveforms for electrocardiography, and experimentally acquired dictionary elements. Overall, our results show that DictXAI provides more interpretable, actionable, and architecture-agnostic insights than classical XAI or existing concept-based approaches.
| Subjects: | Machine Learning (cs.LG); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.10301 [cs.LG] |
| (or arXiv:2610.10301v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10301 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Grégoire Montavon [view email]
[v1]
Wed, 7 Oct 2026 15:56:15 UTC (1,713 KB)
来源:arXiv:cs.LG · arxiv.org