arXiv:cs.AI· Xiqiao Xiong, Moxin Li, Zhixin Ma, Ouxiang Li, Wenjie Wang, Fuli Feng, Xiangnan He·· 5 小时前AI 评分50
HASTE:用稀疏证据演化 Agent harness 以应对新兴攻击
HASTE: Evolving Agent Harnesses Against Emerging Attacks Using Sparse Evidence
AI 导读
中科大等机构研究者在 arXiv 发布论文 HASTE(arXiv:2610.02920),提出一个多智能体框架,从威胁报告、少量攻击示例等稀疏证据出发自动演化 Agent harness 防御。
正文
Abstract:Agent harnesses play a critical role in defenses by enforcing safety constraints to prevent unsafe actions. However, rapidly emerging attacks outpace manual harness adaptation, motivating automated harness evolution. Yet the signals available for harness evolution are often sparse, such as brief descriptions or a few attack examples in threat reports and preprints. To address this limitation, we introduce HASTE, a multi-agent framework that evolves agent harnesses from sparse threat evidence through an adversarial interplay between safety-specification generation and attack-case generation. Safety specifications guide harness updates toward addressing identified safety vulnerabilities, while attack cases probe for remaining safety vulnerabilities after each update. By feeding evaluation outcomes back into both processes, HASTE enables harness evolution against emerging attacks beyond the initially observed evidence. Experimental results across multiple backbone models, attack types, and evidence forms show that HASTE consistently reduces attack success rates while preserving benign-task utility. The code is available at this https URL.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.02920 [cs.AI] |
| (or arXiv:2610.02920v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02920 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xiqiao Xiong [view email]
[v1]
Fri, 2 Oct 2026 07:10:08 UTC (570 KB)
来源:arXiv:cs.AI · arxiv.org