跳到正文
arXiv:cs.LG· Zexin Li, Ruili Yao, Yiming Zeng, Xiaoxue Gao·· 4 小时前AI 评分36

提升针对深度强化学习的可迁移对抗攻击

Boosting Transferable Adversarial Attacks against Deep Reinforcement Learning

AI 导读

针对深度强化学习(DRL)的对抗攻击多假设白盒访问受害者策略,该论文研究基于迁移的黑盒攻击。作者发现直接移植图像分类的可迁移攻击(FGSM、MI-FGSM、NI-FGSM)虽能迁移,但其强度不超过同等预算的随机噪声,随后提出轨迹级攻击,通过环境可微模型与温度平滑代理策略在滚动时域上优化扰动序列。

正文

View PDF HTML (experimental)

Abstract:Most adversarial attacks on deep reinforcement learning (DRL) assume white-box access to the victim policy, which rarely holds in practice. This paper studies transfer-based black-box attacks on DRL: the attacker crafts observation perturbations on a white-box surrogate agent and feeds them to an unknown victim. We formulate the attack as return minimization under a per-step perturbation budget. We first show that transplanting transferable image-classification attacks (FGSM, MI-FGSM, and NI-FGSM) with a per-step objective yields perturbations that transfer but are no stronger than random noise of the same budget. We then propose a trajectory-level attack that optimizes a sequence of perturbations over a receding horizon through a differentiable model of the environment and a temperature-smoothed surrogate policy, with the same optimizers. On CartPole-v1 with ten DQN and DDQN agents and 100 surrogate--victim pairs, the trajectory-level attack outperforms per-step attacks and random noise in the white-box, cross-model, and cross-algorithm settings.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.06083 [cs.LG]
  (or arXiv:2610.06083v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.06083

arXiv-issued DOI via DataCite

Submission history

From: Zexin Li [view email]
[v1] Mon, 5 Oct 2026 10:14:05 UTC (63 KB)
[v2] Wed, 7 Oct 2026 06:36:14 UTC (64 KB)

来源:arXiv:cs.LG · arxiv.org