跳到正文
arXiv:cs.AI· Suxin Ji, Hungtao Wan, Shaoxuan Chen, An Zhang·· 6 小时前AI 评分44

面向 LLM 智能体的语义行为水印:抗改写且防伪造的来源追溯

Semantic Behavioral Watermarking: Paraphrase-Robust and Forgery-Resistant Provenance for LLM Agents

AI 导读

研究者提出语义行为水印 SBW,在历史条件下的语义动作簇上嵌入水印,并用密钥化抗碰撞分桶替代公开簇分桶,在随机预言机模型下分桶分配不可预测。在 ToolBench 和 ALFWorld 上,该方法在改写下的检测率最高达 0.97,并将自适应伪造从 100% 降至假阳性下限。其局限是链式重放仍保持 0.76-0.98,且抗改写的代价是每步水印容量约减半。

正文

View PDF HTML (experimental)

Abstract:Behavioral watermarking embeds an owner identifier in an LLM agent's high-level action choices, giving provenance without touching output tokens. Prior agent watermarks break in two ways. First, all three prior schemes bind the watermark to the exact action symbol, so renaming a tool desynchronizes decoding even when the observation is untouched; in AgentMark's own robustness test, paraphrasing the observation alone drops bit-recovery to 16.8%. Second, every prior agent watermark studies only removal: none asks whether an adversary can forge a trajectory that verifies as someone else's, a question answered affirmatively for text watermarks (Jovanović et al., 2024). We present Semantic Behavioral Watermarking (SBW): watermarking over semantic action clusters under history conditioning, with the public-cluster bin replaced by keyed collision-resistant binning whose fresh-bucket assignment is provably unpredictable in the random-oracle model. Across five agent models (3B-14B, four vendors) and three encoders the ordering holds on both benchmarks: on ToolBench (600 trajectories per model) detection under rewriting is 0.49-0.66 for cluster-level versus 0.05-0.17 for exact-symbol at a permutation-calibrated 1% FPR, at 72-83% choice agreement against 22-27% for logit biasing; on ALFWorld (100 episodes per model) it is 0.92-0.97 versus 0.00-0.01. Keyed binning takes adaptive forgery from 100% to the false-positive floor at the primary operating point (bge, r=64). We also mark the boundary that guarantee does not cover: when the adversary copies the victim's own steps, shuffled splicing is neutralized (0.000 on Qwen2.5-3B) but chained replay remains at 0.76-0.98 across the five models, reported as open. Paraphrase robustness costs about half of the per-step watermark capacity. Code is available at this https URL.
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.08668 [cs.CR]
  (or arXiv:2610.08668v1 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2610.08668

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Suxin Ji [view email]
[v1] Tue, 6 Oct 2026 16:50:42 UTC (327 KB)

来源:arXiv:cs.AI · arxiv.org