跳到正文
arXiv:cs.AI· Yunseok Lee, Yunji Kim, Woojin Lee·· 4 小时前AI 评分56

arXiv 论文提出 ICoA:针对工具调用 LLM Agent 的隐蔽间接提示注入攻击

Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents

AI 导读

论文将间接提示注入的攻击成功率拆分为隐蔽成功率(CSR)与显性成功率(OSR),区分用户能否在 Agent 最终回复中察觉注入。分析发现隐蔽与否取决于注入后 Agent 是否把控制权交回用户任务,据此提出 ICoA 攻击,在 AgentDojo 上对四个目标模型取得最高 CSR,比最强基线高 3.79-12.01 个百分点。论文被 EMNLP 2026 主会接收为 Oral。

正文

View PDF HTML (experimental)

Abstract:As LLM agents take real-world actions through tools, indirect prompt injection (IPI) has emerged as a serious threat. The standard metric, Attack Success Rate (ASR), counts whether an injection succeeds but ignores what the user notices in the agent's final response. Looking at successful injection traces, we find two distinct outcomes: the agent executes the injection while returning an otherwise normal response, or reports the injected action in its final response, giving the user a chance to notice. We call these covert and overt successes. From the user's perspective, we decompose ASR into the Covert Success Rate (CSR), counting successes leaving no trace in the final response, and the Overt Success Rate (OSR), counting successes the user can detect. To understand what drives the gap, we analyze successful trajectories and find that the agent's behavior after the injection separates covert from overt: covert traces hand control back to the user task before ending, while overt traces end at the attack itself. This split follows from the ReAct format, where the final response summarizes the most recent action. Building on this observation, we propose ICoA (Induced Covert Attack), an IPI attack designed to induce covert outcomes by steering the agent back to the user task after executing the injection. Across four target models on AgentDojo, ICoA achieves the highest CSR, with gains of 3.79-12.01 percentage points over the strongest baseline.
Comments: EMNLP 2026 Main (Oral), Project website: this https URL
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2608.30362 [cs.AI]
  (or arXiv:2608.30362v3 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2608.30362

arXiv-issued DOI via DataCite

Submission history

From: Yunseok Lee [view email]
[v1] Mon, 31 Aug 2026 07:20:45 UTC (2,525 KB)
[v2] Tue, 1 Sep 2026 02:14:01 UTC (2,525 KB)
[v3] Fri, 2 Oct 2026 08:20:06 UTC (2,525 KB)

来源:arXiv:cs.AI · arxiv.org