arXiv:cs.LG· Yi Wang, Baicheng Chen, Yu Wang, Jian Zhao, Yilei Chen, Tianxing He·· 3 小时前AI 评分40
RMCW:基于 Reed-Muller 码的抗删除语言模型水印
RMCW: A Deletion-Robust Watermark Based on Reed--Muller Codes for Language Models
AI 导读
研究者提出 RMCW,一种基于 Reed-Muller 码的 LLM 水印方法,通过搜索存活的局部代数结构而非全局码字恢复来抵御删除攻击。在 C4 和 ELI5 数据集上,使用 OPT-1.3B 与 Llama-3.1-8B-Instruct 的实验显示,RMCW 保持强干净文本可检测性,并在多种删除与改写攻击下优于或持平基线方法。代码已开源。
正文
Abstract:Large Language Model (LLM) watermarking provides a lightweight mechanism for identifying text generated by a specific model, but its robustness remains fragile under post-processing attacks. Deletion attacks are particularly challenging because they shift token positions and break the alignment between observed tokens and their original watermark positions. We propose Reed--Muller Code Watermarking (RMCW), an LLM watermarking method based on Reed--Muller codes. In contrast to global codeword recovery, RMCW searches for surviving local algebraic structure, leveraging the Reed--Solomon consistency induced by affine-line restrictions of Reed--Muller codewords. During generation, RMCW injects a Reed--Muller structure into the sequence via a secret-keyed vocabulary partition. During detection, it maps the given text to keyed vocabulary bins and tests local subsequences for low-degree Reed--Solomon consistency using Berlekamp--Welch tests. Experiments on C4 and ELI5 datasets with OPT-1.3B and Llama-3.1-8B-Instruct show that RMCW preserves strong clean-text detectability and outperforms or matches the baseline methods under several deletion and rewriting attacks. Our code is available at this https URL.
| Subjects: | Cryptography and Security (cs.CR); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.02817 [cs.CR] |
| (or arXiv:2610.02817v1 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02817 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Baicheng Chen [view email]
[v1]
Fri, 2 Oct 2026 05:05:06 UTC (226 KB)
来源:arXiv:cs.LG · arxiv.org