arXiv:cs.LG· Haoyu Liu, Dingcheng Li, Lukas Rutishauser, Zeyu Zheng·· 7 小时前AI 评分44
DMAST:让多模态 Web 智能体抵御跨模态攻击的双模态多阶段对抗安全训练
Dual-Modality Multi-Stage Adversarial Safety Training: Robustifying Multimodal Web Agents Against Cross-Modal Attacks
AI 导读
针对多模态 Web 智能体双流架构易被 DOM 注入同时污染截图与无障碍树的问题,研究者提出双模态多阶段对抗安全训练框架 DMAST,将智能体与攻击者建模为双人一般和马尔可夫博弈并三阶段共同训练。在分布外任务上,DMAST 将攻击成功率从 41.2% 降至 21.4%,任务完成率相对提升超 60%(6.2%→10.2%),优于现有基于训练的防御并可与提示词防御互补,代码已开源。
正文
Abstract:Multimodal web agents that process both screenshots and accessibility trees are increasingly deployed to interact with web interfaces, yet their dual-stream architecture opens an underexplored attack surface: an adversary who injects content into the webpage DOM simultaneously corrupts both observation channels with a consistent deceptive narrative. Our vulnerability analysis on MiniWob++ reveals that attacks including a visual component far outperform text-only injections, exposing critical gaps in text-centric VLM safety training. Motivated by this finding, we propose Dual-Modality Multi-Stage Adversarial Safety Training (DMAST), a framework that formalizes the agent-attacker interaction as a two-player general-sum Markov game and co-trains both players through a three-stage pipeline: (1) imitation learning from a strong teacher model, (2) oracle-guided supervised fine-tuning that uses a novel zero-acknowledgment strategy to instill task-focused reasoning under adversarial noise, and (3) adversarial reinforcement learning via Group Relative Policy Optimization (GRPO) self-play. On out-of-distribution tasks, DMAST nearly halves the attack success rate (41.2\%$\rightarrow$21.4\%) while raising task completion by over 60\% relative (6.2\%$\rightarrow$10.2\%). Our approach outperforms established training-based defenses and complements prompt-based defenses, demonstrating genuine co-evolutionary progress and robust generalization to complex, unseen environments. Code is available at this https URL.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2603.04364 [cs.LG] |
| (or arXiv:2603.04364v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2603.04364 arXiv-issued DOI via DataCite |
Submission history
From: Haoyu Liu [view email]
[v1]
Wed, 4 Mar 2026 18:29:54 UTC (954 KB)
[v2]
Mon, 5 Oct 2026 23:01:49 UTC (914 KB)
来源:arXiv:cs.LG · arxiv.org