arXiv:cs.LG· Wenhan Chang, Tianqing Zhu, Heng Xu, Wenjian Liu, Wanlei Zhou·· 3 小时前
基于概念推理与数据投毒的复杂数据类机器遗忘框架
Class Machine Unlearning for Complex Data via Concepts Inference and Data Poisoning
AI 导读
研究者提出一种概念引导的投毒式机器遗忘框架,通过显式识别对目标类别或知识贡献最大的概念来指导模型更新。针对图像分类,该方法用 Post-hoc Concept Bottleneck Model 定位表达相关概念的图像区域并构造替换式投毒样本;针对 LLM,则通过多问题诱导目标知识、聚合 Integrated Gradients 识别关键内容并掩码构造投毒训练目标。
正文
Abstract:Machine unlearning aims to remove the influence of specified training data or knowledge from a trained model without requiring full retraining. This capability is particularly important for modern image classifiers and large language models (LLMs), where retraining can be computationally expensive. However, machine unlearning on complex data remains difficult because the target information is often distributed across multiple semantic elements. Existing methods mainly remove samples, modify labels, or edit model parameters to reduce the influence of the forgetting target. These approaches usually do not explicitly identify which semantic concepts connect the forgetting target to the model's prediction or generated response. As a result, it is difficult to determine which information to modify. This uncertainty may leave residual target information or unnecessarily affect knowledge that should be retained. To address this gap, we propose a concept-guided poisoning unlearning framework that explicitly identifies the concepts that contribute most strongly to the target class or knowledge and uses them to guide the model update. For image classification, our method first identifies class-relevant concepts with a Post-hoc Concept Bottleneck Model, localizes the image regions that express these concepts, and constructs replacement-based poisoned samples. For LLMs, it elicits the target knowledge through multiple questions, aggregates Integrated Gradients across the resulting responses to identify consistently important content, and masks this content to construct poisoned training targets. Experiments on multiple datasets show that the proposed framework achieves effective unlearning across different tasks while largely preserving retained model utility.
| Comments: | 18 pages, 11 figures |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2405.15662 [cs.LG] |
| (or arXiv:2405.15662v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2405.15662 arXiv-issued DOI via DataCite |
Submission history
From: Wenhan Chang [view email]
[v1]
Fri, 24 May 2024 15:59:17 UTC (2,503 KB)
[v2]
Wed, 7 Oct 2026 18:18:28 UTC (6,009 KB)
来源:arXiv:cs.LG · arxiv.org