arXiv:cs.LG· Vikhyath Kothamasu, Virginia Smith, Chhavi Yadav·· 6 小时前AI 评分52
DeCompBench:评估智能体面对分解攻击的安全性基准
Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH
AI 导读
研究者发布 DeCompBench,一个专门评估 LLM 智能体在分解攻击下安全性的基准,通过图形化框架按分解式设计原则把有害任务拆成单独无害、可执行的子任务。实验显示,SOTA 智能体对整体有害任务拒绝率高,但对拆分后的变体拒绝率明显更低,且常在无意中完成对抗目标。数据集已公开。
正文
Abstract:LLM-based Agents are becoming increasingly capable and widely deployed, creating growing incentives for adversarial misuse in the real-world. A key emerging threat is Decomposition Attacks \cite{glukhov2024breach, jones2024adversaries} in which a harmful task is broken into simpler, benign subtasks that evade safety mechanisms when executed separately but cumulatively fulfill the malicious intent. Although recent benchmarks assess agent safety in multi-turn and multi-tool-use settings, they do not explicitly capture this form of decompositional misuse and may not represent realistic adversarial execution flows. To this end, we introduce DeCompBench, a benchmark designed specifically to evaluate agentic safety under decomposition attacks. DeCompBench is created with a decomposition-by-design principle using a graphical framework and enables harmful task decomposition into individually benign and executable subtasks with realistic workflows. Our experiments using a custom decomposer show that state-of-the-art agents exhibit high refusal rates on monolithic harmful tasks, but significantly lower refusal rates on their decomposed variants, while often inadvertently fulfilling the adversarial objectives. These findings underscore the need for safety evaluations against decomposition attacks and corresponding defenses. Our dataset is publicly available and can be found at this https URL.
| Subjects: | Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2606.13994 [cs.CR] |
| (or arXiv:2606.13994v2 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2606.13994 arXiv-issued DOI via DataCite |
Submission history
From: Chhavi Yadav [view email]
[v1]
Fri, 12 Jun 2026 00:30:29 UTC (1,080 KB)
[v2]
Tue, 6 Oct 2026 18:31:15 UTC (2,231 KB)
来源:arXiv:cs.LG · arxiv.org