arXiv:cs.CL· Xunlei Chen, Qinghui Gong, Jingkun Xue, Qihe Liu, Shijie Zhou, Fei Ye·· 4 小时前
PTP-U:大语言模型遗忘中的可恢复漂移
Not Every Change Is Necessary: Recoverable Drift in Large Language Model Unlearning
AI 导读
研究人员提出 Propose-Then-Project Unlearning(PTP-U)框架,通过先施加局部解析编辑削弱目标知识关联,再将非目标输出分布与原模型对齐,在满足遗忘约束的同时恢复非目标能力。在三个基准上,PTP-U 实现 81.22%-91.03% 的遗忘率,平均保留 94.20% 的非目标效用,在同等遗忘水平下持续保持更高的非目标效用。
正文
Abstract:Machine unlearning in large language models aims to remove unwanted knowledge while preserving the model's remaining capabilities. Although existing methods use retention objectives or restrict where edits occur, achieving the desired forgetting level can still leave collateral changes that impair non-target behavior. Our recovery comparisons suggest that some of these changes can be reversed while preserving observed forgetting performance. In this work, we present Propose-Then-Project Unlearning (PTP-U), a framework that combines targeted forgetting with the recovery of non-target capabilities. PTP-U first applies local analytic edits to weaken target knowledge associations, then aligns non-target output distributions with those of the original model to recover capabilities while maintaining fixed forgetting constraints. Both stages serve a common goal: satisfying the forgetting requirements while preserving fluent generation and performance on non-target tasks. Across three benchmarks, PTP-U achieves the strongest forgetting-retention trade-off among evaluated methods, reaching 81.22%-91.03% forgetting while preserving 94.20% non-target utility on average. At matched forgetting, PTP-U consistently retains higher non-target utility.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.11915 [cs.CL] |
| (or arXiv:2610.11915v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11915 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xunlei Chen [view email]
[v1]
Thu, 8 Oct 2026 13:12:40 UTC (4,829 KB)
来源:arXiv:cs.CL · arxiv.org