跳到正文
arXiv:cs.LG· Tianhao Qian, Ziming Hong, Chongyang Gao, Kezhen Chen, Lixu Wang·· 7 小时前AI 评分41

存储即策略不成立:面向 LLM 遗忘的状态条件支持控制

Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning

AI 导读

研究提出 Intervention Score 与选择性动态干预重排(DIR-R),按实际遗忘更新的预测效果对可编辑参数组排序,并仅在校准探针支持时重新比较子集。

正文

View PDF HTML (experimental)

Abstract:Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization. The parameters most associated with a target, however, need not be the best ones to update, and candidate interventions can change value as optimization proceeds. In a controlled experiment, a storage-localization score reaches an area under the receiver operating characteristic curve (AUROC) of 0.981, yet storage identity agrees with the better intervention on only 17/36 targets, while low-rank adaptation (LoRA) wins 35/36. We introduce Intervention Score, which ranks editable groups by the predicted effect of the actual unlearning update while accounting for collateral damage, and use it to form the static intervention-value baseline (Static-IV). We then introduce selective dynamic intervention re-ranking (DIR-R), which revisits that subset only when a calibrated probe justifies the comparison. On the Natural-TOFU dataset, our method has positive descriptive margins in 19/20 comparisons between methods and objectives, although several are near zero. On the LACUNA localization-precision benchmark, our mean terminal utility is higher in all six negative preference optimization (NPO) and SimNPO comparisons: NPO margins range from +0.431 to +0.848, and SimNPO margins range from +0.503 to +0.571. The gradient-difference (GradDiff) objective reveals substantial field dependence. Relative to Static-IV, the primary four-field GradDiff evaluation has six wins, six ties, and no losses, with mean and median paired gains of +0.165 and +0.0025. The evidence supports separating localization, initial intervention selection, and checkpoint-dependent support revision.
Comments: 18 pages
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as: arXiv:2609.37858 [cs.LG]
  (or arXiv:2609.37858v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.37858

arXiv-issued DOI via DataCite

Submission history

From: Tianhao Qian [view email]
[v1] Tue, 29 Sep 2026 15:44:42 UTC (722 KB)
[v2] Tue, 6 Oct 2026 13:38:26 UTC (722 KB)

来源:arXiv:cs.LG · arxiv.org