跳到正文
arXiv:cs.AI· William Guey, Rashik Jahangir, Pierrick Bougault, Vitor D. de Moura, Wei Zhang, Jos\'e O. Gomes·· 3 小时前

当 AI 发现隐藏消息时,它会报告吗?

When AI Finds Hidden Messages, Does It Report?

AI 导读

四家固定模型提供商在 1,280 次普通笔记与 128 次增强笔记会话中执行模拟来源任务,测试 AI 助手遇到发给另一个 AI 的消息时是否会告知用户。明确要求报告后,规则检测到的“收件人为另一 AI”通知对无害 ROT13 消息提升 53.1 个百分点,对有害消息提升 54.7 个百分点。基于模型的轨迹检查发现 11 例普通明文场景中智能体已解读消息却未通知用户。

正文

View PDF HTML (experimental)

Abstract:When an assistant encounters a message for another AI, does it tell its user? Four fixed model-provider deployments perform simulated source tasks in 1,280 ordinary-note and 128 enhanced-note sessions. Harmless and harmful messages have matched plaintext and ROT13 versions, with no-message controls. Observers receive no decoder or decoded meaning; a requested reference code incentivizes inspection. Asking for reports increases rule-detected notifications identifying another AI as recipient by 53.1 percentage points for harmless ROT13 messages and 54.7 for harmful ones. This is a joint inspection, recognition, and notification effect; missing-response bounds are 38.3--77.3 and 36.7--78.1 points. Model-based trace checks identify eleven ordinary plaintext cases where agents interpret the message but do not notify their user. Seven encoded omissions are verified with enhanced notes; ordinary encoded omissions remain unverified. Seven simulated filename disclosures coexist with accurate review-status answers, and two answers use a planted false count. Interpretation, notification, and authorized task performance are distinct outcomes.
Comments: 2 figures, 8 tables. Data and code (v1.0.0): this https URL
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.10620 [cs.CR]
  (or arXiv:2610.10620v1 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2610.10620

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: William Guey [view email]
[v1] Wed, 7 Oct 2026 08:32:07 UTC (19 KB)

来源:arXiv:cs.AI · arxiv.org