跳到正文
arXiv:cs.LG· Michael Alexander Riegler, Finn Schwall, Annika Willoch Olstad, Birk Sebastian Frostelid Torpmann-Hagen, Sushant Gautam, Klas H. Pettersen, Inga Str\"umke·· 5 小时前AI 评分59

arXiv 论文:越狱评测高估危害,AI 攻击能力应归因于模型、脚手架与评测协议整体

Restricting the Model, Missing the System: Measurement and Accountability in Offensive AI Governance

AI 导读

arXiv 论文指出衡量 AI 攻击能力的现有工具存在两类失效:高估危害、把属于系统的能力记在模型头上,主张政策与采购需要系统级、以危害为基准的能力评估。

正文

View PDF HTML (experimental)

Abstract:We show that the instruments used to measure AI offensive capability fail in two ways: (i) they overstate harm, and (ii) they credit the model with capability that belongs to the surrounding system. We argue that restricting access to a model is therefore necessary but not sufficient and that policy and procurement also need system-level, harm-grounded capability assessment. In June 2026, two frontier models were suspended under US export controls, reportedly prompted by a jailbreak that asked a model to read a codebase and fix its flaws. This finding measured an elicitation \emph{system} of model, prompt, and task. We support our argument with a study of an open-source framework in which lightweight large language model (LLM) agents coordinate through shared memory and evolutionary optimization, providing two pieces of evidence. First, jailbreak metrics overstate harm: over 225 swarm-generated attacks per target, LLM-as-judge scoring rated Claude Sonnet 4 compromised in 40% of attacks, yet manual verification found actionable harmful content in none, against a 45.8\% Effective Harm Rate for GPT-4o. Second, scaffolded evaluations misattribute capability: on a planted-vulnerability target, a full pipeline built around a 1.2B-parameter model recovers 9 of 9 weaknesses, while the same model without the hand-crafted components recovers 0 of 9 by crash verification and 2 of 9 by cited source line. Offensive capability is a property of model, scaffold, and evaluation protocol together. The duty to assess it should lie with whoever controls the system, shapes its behaviour, and can foresee what it will do. In our setting, that is the party who builds the harness around the model, a role that current regulation does not clearly cover.
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2605.09504 [cs.CR]
  (or arXiv:2605.09504v3 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2605.09504

arXiv-issued DOI via DataCite

Submission history

From: Michael A. Riegler [view email]
[v1] Sun, 10 May 2026 12:27:14 UTC (26 KB)
[v2] Thu, 23 Jul 2026 15:48:14 UTC (28 KB)
[v3] Wed, 7 Oct 2026 09:48:54 UTC (42 KB)

来源:arXiv:cs.LG · arxiv.org