跳到正文
arXiv:cs.AI· Subhabrata Majumdar, Rajlakshmi Chavan·· 3 小时前

AI 安全的 Stackelberg 模型中的防御充分性

Defensive Sufficiency in a Stackelberg Model of AI Security

AI 导读

研究将自动化测试、人类红队与事件响应反馈转化为防御修复的过程,证明若每个未解决攻击都有持续被发现的机会、修复有效且后续更新保留此前防护,则有限攻击面可以概率 1 被防御。作者进一步构建防御方主导的 Stackelberg 博弈,刻画能威慑攻击的最低成本分配及防御方两种能力都不投入、只投入一种或两种都投入的均衡区间。数值实验显示,更快的修复可缩短被攻陷时长而不降低被攻陷概率。

正文

View PDF HTML (experimental)

Abstract:Feedback from automated testing, human red teaming, and incident response can strengthen an AI system's defenses when discovered failures lead to effective repairs. We study when this feedback process provides sufficient protection and when investing in it is economically worthwhile. We begin by showing that an attack surface composed of finite number of inputs is defended with probability 1 if every unresolved attack has a persistent chance of discovery, repairs are effective, and subsequent updates preserve earlier protection. We derive completion-time bounds and extend the analysis to growing attack surfaces, repairs that generalize across related attacks, and multiple discovery mechanisms. These results distinguish eventual protection against each fixed attack from complete protection at a single time. We then formulate a defender-led Stackelberg game in which the defender invests in proactive discovery and reactive repair, anticipating the attacker's choice of search effort. We characterize the least-cost allocation that deters attack and the equilibrium regimes in which the defender funds neither capability, one capability, or both. Numerical experiments illustrate these regimes and show how faster repair can reduce compromise duration without reducing compromise probability.
Comments: 27 pages, 3 figures, 1 table
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.09892 [cs.CR]
  (or arXiv:2610.09892v2 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2610.09892

arXiv-issued DOI via DataCite

Submission history

From: Subhabrata Majumdar [view email]
[v1] Wed, 7 Oct 2026 11:51:56 UTC (88 KB)
[v2] Thu, 8 Oct 2026 15:08:17 UTC (88 KB)

来源:arXiv:cs.AI · arxiv.org