跳到正文
arXiv:cs.AI· Shaswata Mitra, Raj Patel, Subash Neupane, Sudip Mittal, Md Rayhanur Rahman, Shahram Rahimi·· 6 小时前AI 评分47

规则止于何处,裁判始于何处:测量多智能体系统安全中的判断边界

Where Rules End and Judges Begin: Measuring the Judgment Boundary in Multi-Agent Systems Security

AI 导读

研究将 LLM 多智能体系统(MAS)的防御组织为五项原则并实现为 DEFER1,通过 28 项检查级联并交由四名裁判组成的评审团处理剩余请求。跨四个领域的独立测试显示,攻击成功率从约 30.0% 降至约 3.0%,其中 78% 被拦截的攻击由确定性检查完成。在安全运营领域,仅四分之一提案进入裁判环节,但风险评分审批门会错误批准大多数攻击提案而只批准少数合法提案。

正文

View PDF HTML (experimental)

Abstract:LLM-based multi-agent systems (MAS) engage tools, share memory, and delegate tasks, often encountering adversarial content. Current defenses for MAS are typically evaluated in isolation, focusing on one attack type at a time, which can lead to costly and hard-to-audit outcomes. This study organizes defenses into five principles, implementing them as DEFER1 (DEterministic-First Enforcement with Residual judgment), which includes a cascade of 28 checks that blocks what it can and refers the rest to a panel of four judges. In independent testing across four domains, attack success rates drop from about 30.0% to approximately 3.0%, with 78% of blocked attacks handled by deterministic checks. Only a quarter of proposals reach the judges in the security-operations domain, illustrating that the rules provide security for attacks violating clear policies, while judges manage those that only misrepresent intent. Both systems have weaknesses, such as a risk-score approval gate that inaccurately approves most attack proposals but few legitimate ones, highlighting the challenges in assessing threats accurately.
Comments: 26 pages, 20 figures, 24 tables
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR); Multiagent Systems (cs.MA)
Cite as: arXiv:2610.07657 [cs.AI]
  (or arXiv:2610.07657v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.07657

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Shaswata Mitra [view email]
[v1] Tue, 6 Oct 2026 02:52:51 UTC (255 KB)

来源:arXiv:cs.AI · arxiv.org