arXiv:cs.AI· Anna Mazhar, Sainyam Galhotra·· 3 小时前
CAFÉ:机器遗忘的因果黑盒测试
CAF\'E: Causal Black-Box Testing of Machine Unlearning
AI 导读
研究人员提出 CAFÉ,一种基于规格的黑盒测试方法,仅凭已部署模型的预测即可检测机器遗忘的残留影响。它干预目标特征并将变化传播至下游特征,同时衡量直接与间接因果路径上的残留影响,并给出细粒度诊断。在两个因果网络基准、四种遗忘方法上,CAFÉ 以 0.92–0.93 的成对准确率对残留影响排序,而现有检查最高仅 0.71,且会双向误判。
正文
Abstract:Machine learning models are increasingly deployed as software components that must evolve as requirements change. When specific training records or features must no longer influence a deployed model, machine unlearning aims to remove that influence without retraining from scratch. Because unlearning is often approximate, its effectiveness must be tested. Such tests must often treat the model as a black box, without access to its parameters, training history, or unlearning procedure. Features pose a further challenge: even after a feature is removed from a model's inputs, its influence can persist through downstream features. Many existing checks examine only the feature's direct use and can therefore certify a model that still depends on it. We frame unlearning testing as specification-based testing and present CAFÉ, which, using only a deployed model's predictions, intervenes on the feature, propagates the change to its downstream features, and checks whether the predictions still respond. CAFÉ measures a target's residual influence through both its direct and indirect causal paths, and its fine-grained diagnostics show which channels and subgroups still carry it. On two causal-network benchmarks with four unlearning methods, CAFÉ ranks residual influence with 0.92--0.93 pairwise accuracy, against at most 0.71 for existing checks, which fail in both directions: they certify models whose influence persists through downstream features and flag correctly unlearned ones. On real census data, CAFÉ likewise exposes influence that survives retraining yet goes unnoticed by direct-input checks.
| Subjects: | Software Engineering (cs.SE); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2509.16525 [cs.SE] |
| (or arXiv:2509.16525v2 [cs.SE] for this version) | |
| https://doi.org/10.48550/arXiv.2509.16525 arXiv-issued DOI via DataCite |
Submission history
From: Anna Mazhar [view email]
[v1]
Sat, 20 Sep 2025 04:19:37 UTC (666 KB)
[v2]
Thu, 8 Oct 2026 03:18:24 UTC (148 KB)
来源:arXiv:cs.AI · arxiv.org