跳到正文
arXiv:cs.AI· Nektarios Kalampalikis, Kavya Gupta, Georgi Vitanov, Isabel Valera·· 3 小时前

面向合理概念瓶颈模型:CREAM 框架让 CBM 灵活编码概念关系

Towards Reasonable Concept Bottleneck Models

AI 导读

研究者提出 CREAM(Concept REAsoning Models)框架,可让实践者在 Concept Bottleneck Models 的推理中显式编码概念-概念与概念-任务关系,支持互斥、层级关联等任意类型。

正文

View PDF HTML (experimental)

Abstract:We propose a novel, flexible, and efficient framework for designing Concept Bottleneck Models (CBMs) that enables practitioners to explicitly encode and extend their prior knowledge and beliefs about the concept-concept ($C-C$) and concept-task ($C \to Y$) relationships within the model's reasoning when making predictions. The resulting $\textbf{C}$oncept $\textbf{REA}$soning $\textbf{M}$odels (CREAMs) architecturally encode arbitrary types of $C-C$ relationships such as mutual exclusivity, hierarchical associations, and/or correlations, as well as potentially sparse $C \to Y$ relationships. Moreover, CREAM can optionally incorporate a regularized side-channel to complement the potentially {incomplete concept sets}, achieving competitive task performance while encouraging predictions to be concept-grounded. To evaluate CBMs in such settings, we introduce a $C \to Y$ agnostic metric that quantifies interpretability when predictions partially rely on the side-channel. In our experiments, we show that, without additional computational overhead, CREAM models support efficient interventions, can avoid concept leakage, and achieve black-box-level performance under missing concepts. We further analyze how an optional side-channel affects interpretability and intervenability. Importantly, the side-channel enables CBMs to remain effective even in scenarios where only a limited number of concepts are available.
Comments: 34 pages, 22 figures, Updated to the published version
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
Cite as: arXiv:2506.05014 [cs.LG]
  (or arXiv:2506.05014v3 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2506.05014

arXiv-issued DOI via DataCite

Journal reference: Transactions on Machine Learning Research, 2026

Submission history

From: Nektarios Kalampalikis [view email]
[v1] Thu, 5 Jun 2025 13:22:29 UTC (2,578 KB)
[v2] Sat, 11 Apr 2026 16:01:20 UTC (3,323 KB)
[v3] Thu, 8 Oct 2026 14:45:49 UTC (3,292 KB)

来源:arXiv:cs.AI · arxiv.org