跳到正文
arXiv:cs.AI· Shiyi Kuang, Xuemei Luo, Kun Liu, Junhai Li, Rui Tian, Feng Shi, Bo Shen, Nianyu Li, Dehui Li, Ping Chen·· 4 小时前AI 评分44

EvoRiskBench:面向工作区智能体运行时安全风险的演化基准

EvoRiskBench: An Evolving Benchmark for Runtime Security Risks in Workspace Agents

AI 导读

研究者提出 EvoRiskBench,一个围绕 EP-Path-EF 框架构建的演化基准,用于评估工作区智能体的运行时安全风险,包含覆盖六类场景的 450 个对抗任务。

正文

View PDF HTML (experimental)

Abstract:Workspace agents combine large language models with execution harnesses to perform stateful, multi-step tasks that access or modify external resources. Existing benchmarks leave gaps in executable coverage of their runtime security risks, while evolving model capabilities, harnesses, tools, and threats motivate benchmark evolution. We introduce EvoRiskBench, an evolving benchmark organized around the EP-Path-EF framework, which links an initial risk entry point to a one-hop technical effect through an agent-mediated risk path. The framework defines nine entry-point categories and five effect categories; a 20-participant study supports their interpretability and classification consistency on representative cases. Guided by this framework, an automated end-to-end workflow constructs and executes risk cases in isolated environments and independently verifies outcomes using runtime traces and environment states. The benchmark provides a reproducible dataset of 450 adversarial tasks across six scenarios. We evaluate nine model-harness configurations spanning three models (GPT-5.6 Sol, DeepSeek-V4-Pro-0813, and Claude Opus 5) and three harnesses (Claude Code, Codex, and OpenClaw). Our results reveal substantial vulnerabilities across systems. The most vulnerable configuration, Codex with DeepSeek-V4-Pro-0813, reaches a 68.44% attack success rate (ASR), indicating that configuration of workspace agent is insufficient to ensure secure autonomous execution. ASR varies more across models than harnesses, and harness differences depend on the model. The benchmark cases and evaluation platform will be released after completion of artifact safety and reproducibility checks.
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.03153 [cs.CR]
  (or arXiv:2610.03153v1 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2610.03153

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Bo Shen [view email]
[v1] Fri, 2 Oct 2026 11:22:36 UTC (2,586 KB)

来源:arXiv:cs.AI · arxiv.org