跳到正文
arXiv:cs.AI· Mayank Rathee, Alexander Stepanov, Shalin Madabhavi, Jinhao Zhu, Raluca Ada Popa, Ion Stoica·· 4 小时前AI 评分41

Pincer:用数字孪生为 AI 智能体做资源授权

Pincer: Resource Authorization for Agents using a Digital Twin

AI 导读

Pincer 是一种在资源层运行的智能体防御方案,核心是一个隔离上下文的数字孪生模型,能自动学习并执行用户专属的最小权限策略,充当用户对智能体权限请求的代理。研究还提出一个基于多日用户-智能体交互记录的用户中心数据集用于模拟学习阶段。评估显示 Pincer 在安全性和实用性上均优于多种 LLM 评判基线及 Conseca 的改编版本,在部分攻击类型上安全提升显著。

正文

View PDF HTML (experimental)

Abstract:Coding agents have become increasingly long-horizon, autonomous, reliant on general-purpose shell and maintain their own persistent memory for self-improvement. While these capabilities have made the agents powerful, they have also made them harder to defend against external adversaries. Defenses that restrict this architecture --- typed tools, information-flow control, or policy prediction engines --- give up too much functionality to be adopted. Agents deployed today (e.g. Claude, Codex) rely on a combination of user-mediated and automode sandboxing as their primary defense. In user-mediated sandboxing, user-maintained policies decay over time and repeated permission requests cause user fatigue, while auto mode's tool-call classifiers learn no user-specific policy and are not meant to defend against adversarial setups. Pincer is a new defense that operates at the resource layer and works alongside existing defenses at the tool-call layer like the auto mode. At the core of Pincer lies a digital twin, an isolated-context model that automatically learns and enforces dynamic user-specific least-privilege policies. The digital twin keeps continually learning the user's preferences allowing it to act as the user's proxy for the agent's permission requests. To emulate the learning phase, we propose a new usercentric dataset with examples following a multi-day transcript of user-agent interaction. Our evaluation shows that Pincer performs strongly on both security and utility in comparison to several baselines which includes variants of LLM judges and adaptations of Conseca (HotOS '25). We highlight attack types where Pincer's design leads to a significant security improvement compared to all other baselines, while outperforming the baselines even for other types of attacks.
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.02569 [cs.CR]
  (or arXiv:2610.02569v1 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2610.02569

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Mayank Rathee [view email]
[v1] Thu, 1 Oct 2026 23:03:12 UTC (537 KB)

来源:arXiv:cs.AI · arxiv.org