跳到正文
arXiv:cs.LG· Toluwani Aremu, Noor Hussein, Munachiso Nwadike, Samuele Poppi, Jie Zhang, Karthik Nandakumar, Neil Gong, Nils Lukas·· 5 小时前AI 评分43

通过随机密钥选择缓解生成模型中的水印伪造

Mitigating Watermark Forgery in Generative Models via Randomized Key Selection

AI 导读

研究者提出一种随机化水印密钥选择的防御方案,每次查询随机选密钥,仅当恰好一个密钥检测到水印时才判定内容为真,可应用于任意现有水印方法且不进一步降低模型效用。在 r=4 个密钥时,有害文本伪造成功率从单密钥的 87% 降至 1%,图像初步研究从 100% 降至 2%,计算开销可忽略。

正文

View PDF HTML (experimental)

Abstract:Watermarking enables GenAI providers to verify whether content was generated by their models. A watermark is a hidden signal in the content, whose presence can be detected using a secret watermark key. A core security threat are forgery attacks, where adversaries insert the provider's watermark into content \emph{not} produced by the provider, potentially damaging their reputation and undermining trust. Existing defenses resist forgery by embedding many watermarks with multiple keys into the same content, which can degrade model utility. However, forgery remains a threat when attackers can collect sufficiently many watermarked samples. We propose a defense with a sample-count-independent upper bound on forgery success for blind attackers, conditional on key-symmetric, independent detector outcomes. Our scheme does not further degrade model utility. We randomize the watermark key selection for each query and accept content as genuine only if a watermark is detected by \emph{exactly} one key. Unlike cryptographic watermarks that rely on computational hardness assumptions and require designing new watermarking schemes from scratch, our method can be applied to any existing watermarking method to improve its forgery resistance. We focus on text watermarking, but our defense is modality-agnostic, since it treats the underlying watermarking method as a black-box. To show this, we include a preliminary study on image watermarking using Tree-Ring. Separately from this conditional guarantee, we empirically observe that, at $r=4$ keys, harmful-text forgery success drops from as high as $87\%$ with a single key to as low as $1\%$ against the adaptive blind attackers that we evaluate, at negligible computational overhead; a preliminary image study shows a reduction from $100\%$ to $2\%$.
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2507.07871 [cs.CR]
  (or arXiv:2507.07871v5 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2507.07871

arXiv-issued DOI via DataCite

Submission history

From: Toluwani Aremu [view email]
[v1] Thu, 10 Jul 2025 15:52:32 UTC (3,405 KB)
[v2] Sat, 2 Aug 2025 12:28:08 UTC (3,880 KB)
[v3] Sat, 27 Sep 2025 07:12:29 UTC (3,901 KB)
[v4] Mon, 11 May 2026 08:00:17 UTC (3,874 KB)
[v5] Fri, 2 Oct 2026 17:58:26 UTC (3,977 KB)

来源:arXiv:cs.LG · arxiv.org