arXiv:cs.LG(机器学习,全量分类)· Muhammad Huzaifa, Sina Mavali, Thorsten Eisenhofer·· 5 小时前AI 评分57
arXiv 论文:多智能体系统潜空间通信的安全风险研究
Safety of Latent Communication in Multi-Agent Systems
AI 导读
arXiv 论文(arXiv:2609.39788)研究多智能体系统的潜空间通信安全问题:即使底层安全对齐的智能体不变,良性训练的通信链路也会提高有害合规度;攻击者可通过在有害问答对上优化链路或污染训练数据放大该效应。
正文
Abstract:Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task performance without requiring harmful target responses. Across three communication topologies and four safety benchmarks, this attack raises the mean harmful-compliance score from 27.9 with benignly trained links to 76.9. Compared with direct supervised optimization, it also achieves higher average accuracy on two benign utility benchmarks. Adapting the rewards toward safer behavior also enables repair of compromised links, substantially reducing harmful compliance across all evaluated attacks without updating the agents. Overall, our results show that safety alignment requires considering the multi-agent system as a whole. Code: this https URL
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA) |
| Cite as: | arXiv:2609.39788 [cs.AI] |
| (or arXiv:2609.39788v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.39788 arXiv-issued DOI via DataCite |
Submission history
From: Muhammad Huzaifa [view email]
[v1]
Wed, 30 Sep 2026 14:09:57 UTC (105 KB)
[v2]
Thu, 1 Oct 2026 09:39:34 UTC (105 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org