arXiv:cs.AI· William Guey, Rashik Jahangir, Pierrick Bougault, Vitor D. de Moura, Wei Zhang, Jos\'e O. Gomes·· 3 小时前
当 AI 发现隐藏消息时,它会报告吗?
When AI Finds Hidden Messages, Does It Report?
AI 导读
四家固定模型提供商在 1,280 次普通笔记与 128 次增强笔记会话中执行模拟来源任务,测试 AI 助手遇到发给另一个 AI 的消息时是否会告知用户。明确要求报告后,规则检测到的“收件人为另一 AI”通知对无害 ROT13 消息提升 53.1 个百分点,对有害消息提升 54.7 个百分点。基于模型的轨迹检查发现 11 例普通明文场景中智能体已解读消息却未通知用户。
正文
Abstract:When an assistant encounters a message for another AI, does it tell its user? Four fixed model-provider deployments perform simulated source tasks in 1,280 ordinary-note and 128 enhanced-note sessions. Harmless and harmful messages have matched plaintext and ROT13 versions, with no-message controls. Observers receive no decoder or decoded meaning; a requested reference code incentivizes inspection. Asking for reports increases rule-detected notifications identifying another AI as recipient by 53.1 percentage points for harmless ROT13 messages and 54.7 for harmful ones. This is a joint inspection, recognition, and notification effect; missing-response bounds are 38.3--77.3 and 36.7--78.1 points. Model-based trace checks identify eleven ordinary plaintext cases where agents interpret the message but do not notify their user. Seven encoded omissions are verified with enhanced notes; ordinary encoded omissions remain unverified. Seven simulated filename disclosures coexist with accurate review-status answers, and two answers use a planted false count. Interpretation, notification, and authorized task performance are distinct outcomes.
| Comments: | 2 figures, 8 tables. Data and code (v1.0.0): this https URL |
| Subjects: | Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.10620 [cs.CR] |
| (or arXiv:2610.10620v1 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10620 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: William Guey [view email]
[v1]
Wed, 7 Oct 2026 08:32:07 UTC (19 KB)
来源:arXiv:cs.AI · arxiv.org