arXiv:cs.LG· Zexin Li, Ruili Yao, Yiming Zeng, Xiaoxue Gao·· 4 小时前AI 评分36
提升针对深度强化学习的可迁移对抗攻击
Boosting Transferable Adversarial Attacks against Deep Reinforcement Learning
AI 导读
针对深度强化学习(DRL)的对抗攻击多假设白盒访问受害者策略,该论文研究基于迁移的黑盒攻击。作者发现直接移植图像分类的可迁移攻击(FGSM、MI-FGSM、NI-FGSM)虽能迁移,但其强度不超过同等预算的随机噪声,随后提出轨迹级攻击,通过环境可微模型与温度平滑代理策略在滚动时域上优化扰动序列。
正文
Abstract:Most adversarial attacks on deep reinforcement learning (DRL) assume white-box access to the victim policy, which rarely holds in practice. This paper studies transfer-based black-box attacks on DRL: the attacker crafts observation perturbations on a white-box surrogate agent and feeds them to an unknown victim. We formulate the attack as return minimization under a per-step perturbation budget. We first show that transplanting transferable image-classification attacks (FGSM, MI-FGSM, and NI-FGSM) with a per-step objective yields perturbations that transfer but are no stronger than random noise of the same budget. We then propose a trajectory-level attack that optimizes a sequence of perturbations over a receding horizon through a differentiable model of the environment and a temperature-smoothed surrogate policy, with the same optimizers. On CartPole-v1 with ten DQN and DDQN agents and 100 surrogate--victim pairs, the trajectory-level attack outperforms per-step attacks and random noise in the white-box, cross-model, and cross-algorithm settings.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.06083 [cs.LG] |
| (or arXiv:2610.06083v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.06083 arXiv-issued DOI via DataCite |
Submission history
From: Zexin Li [view email]
[v1]
Mon, 5 Oct 2026 10:14:05 UTC (63 KB)
[v2]
Wed, 7 Oct 2026 06:36:14 UTC (64 KB)
来源:arXiv:cs.LG · arxiv.org