arXiv:cs.AI(全量分类)· Corinn Tiffany, Wen Zhang, Eugene Bagdasarian, Lillian Tsai·· 5 小时前AI 评分46
Sapien:面向自主 AI 智能体的有状态策略引擎
Sapien: A Stateful Policy Engine for Autonomous AI Agents
AI 导读
Sapien 是一款为自主 AI 智能体强制执行有状态上下文策略的策略引擎,其策略用扩展正则表达式描述允许的工具调用序列,并支持有状态谓词、延迟策略生成与作用域语义检查。它在保持接近无约束智能体效用的同时,即便智能体被完全劫持,也能在 AgentDojo 上排除 93-95% 的攻击、在 Toolathlon 上排除 62-85%,长程任务中拦截量是工具白名单的两倍。
正文
Abstract:Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a policy engine for enforcing stateful contextual policies. A Sapien policy specifies permitted tool-call sequences using a regular expression extended with stateful predicates, deferred policy generation, and scoped semantic checks. We show that Sapien stays within a few percent of an unconstrained agent's utility. Even if the agent is fully hijacked, Sapien's policies rule out 93-95% of attacks on AgentDojo and 62-85% on Toolathlon (twice as many as tool allowlists on long-horizon tasks).
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR) |
| Cite as: | arXiv:2610.00797 [cs.AI] |
| (or arXiv:2610.00797v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00797 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Wen Zhang [view email]
[v1]
Wed, 30 Sep 2026 22:40:00 UTC (134 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org