arXiv:cs.AI· Yukiya Horiba, Koshiro Aoki, Shunsuke Yasuki, Bum Jun Kim, Taiki Miyanishi·· 4 小时前AI 评分34
检测与抑制:针对 VLA 模型对抗补丁的机制性防御
Detect and Suppress: A Mechanistic Defense against Adversarial Patches in VLA Models
AI 导读
研究者用稀疏自编码器(SAE)分析 VLA 模型内部表征,找到一个激活与对抗补丁存在强相关的特征,并仅在线性探针检测到攻击时于推理阶段抑制该特征,无需微调即可提升鲁棒性。在 LIBERO-10 上的评测显示,条件式干预提升了间歇性攻击下的成功率,而持续施加同一干预会显著损害策略性能。
正文
Abstract:Adversarial patches can disrupt Vision-Language-Action (VLA) models by manipulating visual observations, leading to failures in robot control. However, it remains poorly understood which internal mechanisms underlie these failures and how targeted interventions can mitigate them. In this work, we mechanistically analyze VLA representations using a sparse autoencoder (SAE) and identify a feature whose activation strongly correlates with the presence of an adversarial patch. Based on this analysis, we suppress the identified feature at inference time only when a linear probe detects an attack. This intervention improves robustness without the cost of fine-tuning the VLA. We evaluate our method against VLA adversarial patch attacks on LIBERO-10. Conditional intervention improves success rate under intermittent attacks, whereas continuously applying the same intervention substantially degrades policy performance. These results show that attack-related internal representations can provide useful targets for VLA adversarial defense and that controlling when to intervene is important for limiting disruption to nominal policy behavior.
| Subjects: | Robotics (cs.RO); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.03498 [cs.RO] |
| (or arXiv:2610.03498v1 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03498 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yukiya Horiba [view email]
[v1]
Fri, 2 Oct 2026 15:57:04 UTC (845 KB)
来源:arXiv:cs.AI · arxiv.org