跳到正文
arXiv:cs.CL· Toluwani Aremu, Samuele Poppi, Nils Lukas·· 3 小时前AI 评分41

Constitution-Guided Watermarking:按需取舍的文本水印框架

Constitution-Guided Watermarking

AI 导读

研究者提出 Constitution-Guided Watermarking,用自然语言原则让供应商为不同请求选择合适的水印权衡配置。离线阶段由预训练推理智能体结合规则与水印实现迭代调优,部署时由独立监控器识别适用规则并检索对应策略,无需改动服务模型。

正文

View PDF HTML (experimental)

Abstract:Watermarking enables language model providers to identify text generated by their models. However, its desired properties can conflict (\ie~stronger watermark signals can degrade text quality), while designs that resist editing may also facilitate forgery. Providers address these trade-offs by choosing configurations that balance competing objectives or prioritize particular properties. Either approach imposes a shared operating point on requests with different requirements, potentially sacrificing quality where wording preservation matters or robustness where reliable attribution is essential. To allow flexible and adaptable designs, we introduce \emph{Constitution-Guided Watermarking}, a framework that selects request-appropriate trade-offs from provider requirements, listed as natural-language principles. \emph{Offline}, a pretrained reasoning agent examines constitutional rules alongside watermark implementations and iteratively refines rule-specific configurations using empirical feedback. \emph{At deployment}, a separate monitor identifies applicable rules and retrieves the corresponding policy, including watermarking exemptions, without modifying the serving model. Furthermore, our framework supports offline parallel optimization and refinement of rule-specific configurations based on evolving provider requirements without affecting deployment, and binds each deployed configuration to its evaluation evidence, making deployment decisions auditable. In a proof-of-concept evaluation using KGW and a five-rule constitution, our framework selects configurations responsive to provider priorities and improves post-paraphrase detection on robustness-prioritized requests by up to $14$ percentage points over fixed configurations, while matching or exceeding all baselines in aggregate quality and clean detection at a nominal $0.1\%$ false-positive rate.
Comments: Working paper (under review)
Subjects: Cryptography and Security (cs.CR); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2610.09552 [cs.CR]
  (or arXiv:2610.09552v1 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2610.09552

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Toluwani Aremu [view email]
[v1] Wed, 7 Oct 2026 06:50:47 UTC (1,170 KB)

来源:arXiv:cs.CL · arxiv.org