跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Hanwei Zhang, Tianma Hu, Gaojie Jin, Xu Cheng, Ronghui Mu·· 5 小时前AI 评分39

可解释但脆弱?概念瓶颈模型在几何-语义扰动下的鲁棒性

Interpretable but Fragile? Robustness of Concept Bottlenecks under Geometric-Semantic Perturbations

AI 导读

研究提出基于生成器的评估框架,在潜空间连续几何扰动与概念空间离散语义干预两类条件下,对比标准分类器与概念瓶颈模型(CBM)的鲁棒性,并用量化敏感度指标与潜空间、概念空间的随机平滑进行可验证评估。结果表明可解释性并不天然带来鲁棒性,概念瓶颈只是改变了敏感性的分布位置与方式,可解释性与鲁棒性是相互独立的目标。该工作已被 NeurIPS 2026 接收。

正文

View PDF HTML (experimental)

Abstract:Concept Bottleneck Models (CBMs) are designed to provide interpretable intermediate representations, yet how such bottlenecks affect robustness remains unclear, with existing studies reporting mixed and sometimes contradictory findings. We argue that these discrepancies arise from conflating different robustness notions and perturbation regimes, rather than from fundamental disagreements about CBMs themselves. To disentangle these factors, we introduce a generator-based evaluation framework that enables controlled comparisons between standard classifiers and CBMs under two distinct perturbation types: continuous geometric perturbations in latent space and discrete semantic interventions in concept space. Within this framework, we evaluate robustness both empirically, via prediction and concept-level sensitivity metrics, and certifiably, using randomized smoothing in latent and concept spaces. Across experiments, we reconcile previously conflicting findings by clarifying when, and in what sense, concept bottlenecks do or do not improve robustness. By further analyzing robustness under varying task conditions, including class semantic similarity and concept vocabulary size, we show that interpretability does not inherently confer robustness. Instead, concept bottlenecks shift where and how sensitivity manifests, revealing a nuanced interpretability robustness trade off that depends critically on the perturbation regime and task structure. Together, our results show that interpretability and robustness are distinct objectives: interpretable intermediate representations do not uniformly improve robustness, but instead redistribute sensitivity across perturbation spaces and model families.
Comments: accepted by NeurIPS 2026
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.38625 [cs.LG]
  (or arXiv:2609.38625v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.38625

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ronghui Mu [view email]
[v1] Tue, 29 Sep 2026 22:35:54 UTC (1,842 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org