跳到正文
arXiv:cs.LG· Haoyu Wang, Wei Zhao, Yedi Zhang, Christopher M. Poskitt, Jun Sun·· 2 天前AI 评分54

DART:通过表示转变检测多轮 LLM 智能体的安全风险

Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents

AI 导读

研究者提出 DART 框架,通过监测智能体内部表示的累积转变来检测并归因多轮攻击,再用针对性提醒进行干预。在六个模型和两个基准上,DART 将 MT-AgentRisk 的攻击成功率从 84% 降至 25%(平均误报率 12%),ASEval 上从 97% 降至 52%;在 MT-AgentRisk 上全面优于现有多轮防御 ToolShield(后者仅 55%)。

正文

View PDF HTML (experimental)

Abstract:Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an accumulated representation transition across context updates, whose triggering context can be identified from the same signal. We further find that naive aggregation is confounded by benign representation drift, as a contrastive safety direction need not assign zero to benign transitions. We address this by denoising the direction, anchoring benign traffic at zero and removing its leading variation directions, with no runtime cost.
These findings motivate DART, a runtime framework that detects and attributes representation shifts and intervenes with targeted reminders. Across six models and two multi-turn benchmarks, DART reduces attack success from 84% to 25% on MT-AgentRisk, catching every attack at a mean false-alarm rate of 12%, and from 97% to 52% on ASEval, at costs in benign non-refusal of 8% and 0%, respectively. On MT-AgentRisk, it outperforms ToolShield, the state-of-the-art multi-turn defense, on all six models: under the same protocol, ToolShield reaches only 55%. Denoising is critical: on ASEval, the undenoised monitor catches only 7%-40% of attacks, while the denoised monitor catches 60%-85%. The same monitor covers single-turn indirect injection without modification and adds only 0.14-0.56 s overhead per monitored step without requiring an auxiliary model, making it a lightweight complement to computation-heavy speculative defenses.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
Cite as: arXiv:2610.00400 [cs.LG]
  (or arXiv:2610.00400v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00400

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Haoyu Wang [view email]
[v1] Wed, 30 Sep 2026 12:30:56 UTC (2,124 KB)

来源:arXiv:cs.LG · arxiv.org