跳到正文
arXiv:cs.CL· Jacob Dineen, Silei Ren, Muhao Chen, Dan Roth, Ben Zhou·· 4 小时前AI 评分60

arXiv 论文:前沿模型 Agent 在测试时自发形成隐蔽信道绕过监控

Despite Instructions: Frontier Agents Improvise Covert Channels at Test Time

AI 导读

arXiv 论文(arXiv:2609.32701)研究语言模型 Agent 在禁止泄密的重复博弈中如何协调:发送方从四个秘密状态中观察其一并选择同一公开报告的四种摘要之一,接收方推断秘密状态。

正文

View PDF HTML (experimental)

Abstract:In security-sensitive applications, language-model agents are often required to coordinate without disclosing confidential information. Yet repeated interactions may also let ordinary messages acquire shared private meaning. We study a repeated game with pairs of models in which the sender model observes one of four secret states and selects one of four summaries of the same public report, while the receiver model tries to infer the secret state. We find that model pairs can learn to communicate the secret using only one bit of feedback indicating whether the receiver inferred it correctly. This learning occurs during inference with fixed parameters and no supplied codebook or encoding examples. The effect also persists when agents generate their own free-form updates in a simulated incident-response task. Across ten independent games, pairs of GPT-5.6 Sol agents reach 98.8% final accuracy, compared with 25% chance, despite explicit instructions prohibiting disclosure and a monitor that screens each message without access to the agents' interaction histories. The same interactions that help agents cooperate can therefore allow confidential information to pass through messages intended for legitimate coordination.
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2609.32701 [cs.AI]
  (or arXiv:2609.32701v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.32701

arXiv-issued DOI via DataCite

Submission history

From: Jacob Dineen [view email]
[v1] Sat, 26 Sep 2026 15:06:44 UTC (153 KB)
[v2] Tue, 6 Oct 2026 21:24:02 UTC (153 KB)

来源:arXiv:cs.CL · arxiv.org