跳到正文
原文
arXiv:cs.AI(全量分类)· Corinn Tiffany, Wen Zhang, Eugene Bagdasarian, Lillian Tsai·· 5 小时前AI 评分46

Sapien:面向自主 AI 智能体的有状态策略引擎

Sapien: A Stateful Policy Engine for Autonomous AI Agents

AI 导读

Sapien 是一款为自主 AI 智能体强制执行有状态上下文策略的策略引擎,其策略用扩展正则表达式描述允许的工具调用序列,并支持有状态谓词、延迟策略生成与作用域语义检查。它在保持接近无约束智能体效用的同时,即便智能体被完全劫持,也能在 AgentDojo 上排除 93-95% 的攻击、在 Toolathlon 上排除 62-85%,长程任务中拦截量是工具白名单的两倍。

正文

View PDF HTML (experimental)

Abstract:Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a policy engine for enforcing stateful contextual policies. A Sapien policy specifies permitted tool-call sequences using a regular expression extended with stateful predicates, deferred policy generation, and scoped semantic checks. We show that Sapien stays within a few percent of an unconstrained agent's utility. Even if the agent is fully hijacked, Sapien's policies rule out 93-95% of attacks on AgentDojo and 62-85% on Toolathlon (twice as many as tool allowlists on long-horizon tasks).
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR)
Cite as: arXiv:2610.00797 [cs.AI]
  (or arXiv:2610.00797v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.00797

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Wen Zhang [view email]
[v1] Wed, 30 Sep 2026 22:40:00 UTC (134 KB)

来源:arXiv:cs.AI(全量分类) · arxiv.org