跳到正文
arXiv:cs.AI· Xinran Zheng, Xin Fan Guo, Zhiqiang Hao, Fan Yang, Xingzhi Qian, Jiawei Du, Jinfeng Xu, Zheng Xing, Shuo Yang, Xingjun Wang·· 6 小时前AI 评分62

APEX:在执行边界主动防护 LLM 智能体的间接提示词注入

APEX: Active Protection at Execution Boundaries for LLM Agents

AI 导读

论文提出 APEX,一种针对 LLM 智能体间接提示词注入(IPI)的主动防御方法,将防护收敛到智能体把内部状态转为外部动作或输出的执行边界。

正文

View PDF HTML (experimental)

Abstract:Indirect prompt injection (IPI) hides adversarial instructions in content that large language model (LLM) agents read at runtime. As agents compose heterogeneous capability units, including Tools, MCP servers, and Skills, the carriers of injection multiply, and defenses built to recognize attack patterns fall behind them. We instead shift defense from covering attack patterns to one stable point: whatever the carrier and however the injection propagates, harm materializes only at the \emph{execution boundary}, where the agent turns internal state into an external action or released output. Safety there turns on two conditions, both settled by the trusted task rather than by the run: whether the proposed effect is authorized, and whether the runtime information reaching it is endorsed by that task. We present APEX, an active defense that enforces both at this boundary from a single authorization contract compiled before untrusted execution: \emph{evidence-gated prevention} admits an effect only when the contract justifies it, while \emph{deception-based exposure} makes unendorsed use reveal itself before the effect commits. Protection therefore follows from what the task permits rather than from how an attack is built, and applies uniformly across capability units without attack-specific policies or taint tracking. Against 13 baselines, APEX attains 0\% attack success on five of six benchmarks and 0.56\% on the sixth, holds 0\% under adaptive attacks on all three capability-unit types, and remains effective across defender backbones. Code is available at this https URL.
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.06966 [cs.CR]
  (or arXiv:2610.06966v1 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2610.06966

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xinran Zheng [view email]
[v1] Sat, 3 Oct 2026 17:53:29 UTC (3,191 KB)

来源:arXiv:cs.AI · arxiv.org