跳到正文
arXiv:cs.CL· Mohamed Dhouib, Clement Elliker, Alexi Canesse, Ma\"el Jenny, Lucas-Andrei Thil, Mahammed El Sharkawy, Sonia Vanier, Elie Bursztein·· 4 小时前AI 评分41

RAISED:用自蒸馏提升 LLM 智能体对提示词注入的鲁棒性

RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents

AI 导读

针对工具调用型 LLM 智能体易受间接提示词注入攻击、且现有训练期防御会损害通用能力的问题,研究者提出训练框架 RAISED,结合自生成与自蒸馏,让模型先在自造的工具使用场景中生成轨迹,再通过自蒸馏令学生在干净与注入版本的同一轨迹上对齐教师的干净上下文行为。该方法大幅降低工具响应中提示词注入的攻击成功率,同时在智能体与通用基准上保持原有能力,优于此前的训练期防御方案。

正文

View PDF HTML (experimental)

Abstract:Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial drift in the model's output distribution, altering its behavior even in benign settings and providing a potential mechanism for utility degradation. We further identify a failure mode of these defenses: On benign tool-use tasks, the model refrains from a step needed to finish an authorized task, particularly when that step is indicated by a tool output. To address these limitations, we introduce RAISED (Robust Attack Invariance through Self-Distillation), a training framework that combines self-generation and self-distillation. The model first generates its own tool-use scenarios, with an emphasis on cases where task completion requires acting on legitimate guidance from tool outputs. Then, through self-distillation, the student is trained to match the teacher's clean-context behavior on both clean and injected variants of the same trajectory. RAISED substantially reduces the attack success rate of prompt injections in tool responses while, unlike prior training-based defenses, preserving utility on both agentic and general-purpose benchmarks.
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2610.06401 [cs.CR]
  (or arXiv:2610.06401v2 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2610.06401

arXiv-issued DOI via DataCite

Submission history

From: Alexi Canesse [view email]
[v1] Mon, 5 Oct 2026 14:21:30 UTC (132 KB)
[v2] Wed, 7 Oct 2026 12:59:03 UTC (132 KB)

来源:arXiv:cs.CL · arxiv.org