跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Sidney Bender, Benedikt Kunz, Ahmed Zeid, Shinichi Nakajima, Klaus-Robert M\"uller, Marco Morik·· 14 小时前AI 评分41

DiDAE:面向视觉基础模型的快速解耦反事实解释方法

Towards Fast and Disentangled Counterfactuals for Visual Foundation Models

AI 导读

研究者提出 DiDAE(Disentangled Diffusion Autoencoders),将冻结的基础模型包裹在条件扩散解码器中,通过解耦字典方向的闭式编辑生成反事实,无需梯度,速度比 SOTA 快至多 2000 倍。

正文

View PDF HTML (experimental)

Abstract:Foundation models remain vulnerable to spurious correlations and ``Clever Hans'' strategies. Explainable machine learning can find and remove such strategies for classifiers without metadata. For foundation models, no such option exists yet. We propose Disentangled Diffusion Autoencoders (DiDAE). DiDAE wraps a frozen foundation model in a conditional diffusion decoder. A counterfactual is one closed-form edit along a direction of a disentangled dictionary, followed by decoding. The dictionary can be supervised (Procrustes) or unsupervised (Singular Value Decomposition, Sparse Autoencoders). No gradients are needed, so DiDAE is up to 2000 times faster than the state of the art. We evaluate on six datasets, two synthetic and four real-world. In a desiderata-driven benchmark on three of them, its counterfactuals are on par with or better than the state of the art, and they repair downstream classifiers through Counterfactual Knowledge Distillation (CFKD), where they beat metadata-based correction. The same machinery can rank a pretrained dictionary against a trained classifier. It returns the few directions the classifier actually reads, each causally verified by a counterfactual that flips the decision, and repairs the classifier along those a teacher marks spurious. The workflow is plug-and-play in our open-source Peal library we publish alongside the paper. With a public dictionary and a pretrained decoder, all that remains is a cheap linear distillation of the classifier and its own fine-tuning.
Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2610.00895 [cs.LG]
  (or arXiv:2610.00895v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00895

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sidney Bender [view email]
[v1] Thu, 1 Oct 2026 01:19:00 UTC (30,672 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org