跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Si Qi Goh, Cap Dang Xuan Kiet, Tat-Jen Cham, Kwok-Yan Lam·· 14 小时前AI 评分33

SIEVE:面向视觉语言模型选择性遗忘的注意力值抑制框架

SIEVE: Selective attention-value Suppression for Vision-Language Models Unlearning

AI 导读

SIEVE 是一个面向视觉语言模型(VLM)的选择性遗忘框架,通过将遗忘样本的注意力值抑制至常数零、并将保留样本表示对齐冻结参考模型,实现针对个人身份信息(PII)的定向遗忘。实验显示,SIEVE 在多种模型-模态设置下取得 SOTA 遗忘性能,同时保持有竞争力的保留效用。消融研究表明,值抑制与负交叉熵提供互补的遗忘信号,而基于参考的值匹配显著降低效用退化。

正文

View PDF HTML (experimental)

Abstract:The ability of vision-language models (VLMs) to associate visual identities with biographical information creates a need for selective unlearning of personally identifiable information (PII) while preserving permitted knowledge about the same individual. This setting is challenging because both sensitive and retained information can share the same visual inputs and intermediate representations. We introduce SIEVE, a simple and effective framework for selective VLM unlearning. SIEVE directly regularizes attention-value representations while also controlling model outputs. SIEVE suppresses attention values for forget examples toward a constant zero, while preserving retain-example representations by matching them to a frozen reference model. These objectives are combined with sequence-level forget and retain supervision, enabling targeted forgetting without largely affecting retained knowledge. Extensive experiments show that SIEVE achieves state-of-the-art performance on unlearning with multiple model-modality settings, while maintaining competitive retained utility. Ablation studies further show that value suppression and negative cross-entropy contribute complementary forgetting signals, while reference-based value matching substantially reduces utility degradation. These results demonstrate that attention values provide an effective intervention point for selective multimodal unlearning when sensitive and retained knowledge are closely related.
Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2610.01962 [cs.LG]
  (or arXiv:2610.01962v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01962

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Si Qi Goh [view email]
[v1] Thu, 1 Oct 2026 16:16:53 UTC (21,221 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org