跳到正文
arXiv:cs.AI· Anna Yoo Jeong Ha, Ronik Bhaskar, Haitao Zheng, Ben Y. Zhao·· 4 小时前AI 评分43

Whiteout:用混淆样本阻止 LLM 泄露个人敏感信息

Mitigating Private Data Leakage in LLMs with Whiteout

AI 导读

研究者提出 Whiteout,一种在个人提出请求后,用精确设计的混淆样本覆写真实个人敏感信息(PSI),从而阻止 LLM 复述生日、电话、住址等内容的工具。在包括一款广泛使用的 OpenAI 模型在内的多种规模与厂商的 LLM 上,Whiteout 能有效阻止目标 PSI 泄露,对模型效用与安全性的影响可忽略,并优于现有方案。该工具还通过了越狱等黑盒攻击以及重学习、量化等白盒自适应攻击的测试。

正文

View PDF HTML (experimental)

Abstract:Modern large language models (LLMs) are trained on massive, largely unfiltered datasets, including content scraped from nearly every accessible website and user inputs. As a result, LLMs often memorize and reproduce personally sensitive information (PSI) such as birth dates, phone numbers, and home addresses. This leads to significant privacy risks, particularly for high-profile individuals such as executives, politicians, and judges. Existing mitigations largely rely on machine unlearning. However, these methods often remove more information than needed, degrade model utility and safety, and are highly vulnerable to attacks.
This paper presents Whiteout, a practical tool that, upon requests by individuals, prevents LLMs from regurgitating their genuine PSIs, by overwriting them using precise and carefully designed obfuscation samples. We evaluate Whiteout on modern LLMs of varying sizes and makers, including a widely-used OpenAI model. Results show that Whiteout effectively prevents disclosure of the targeted PSIs, has negligible impact on model utility and safety, and outperforms existing alternatives. We also test Whiteout against a wide range of countermeasures, from black-box attacks like jailbreaking to white-box adaptive attacks like relearning and quantization. Finally, we conclude with a discussion on the security and ethical implications of Whiteout.
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.02418 [cs.CR]
  (or arXiv:2610.02418v1 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2610.02418

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Anna Yoo Jeong Ha [view email]
[v1] Thu, 1 Oct 2026 19:40:38 UTC (1,325 KB)

来源:arXiv:cs.AI · arxiv.org