arXiv:cs.LG· Jialiang Fan, Shixiong Jiang, Mengyu Liu, Fanxin Kong·· 5 小时前AI 评分48
面向机器人系统的黑盒安全强化学习控制器演示引导观测攻击
Demonstration-Guided Observation Attacks on Black-Box Safe Reinforcement Learning Controllers for Robotic Systems
AI 导读
研究者提出一种演示引导的观测攻击方法,仅凭演示数据即可分析未知安全强化学习控制器:通过逆约束强化学习恢复状态约束与代理策略,并从演示转移中学习动力学,组合梯度生成有界观测扰动,无需受害者参数、梯度或查询。
正文
Abstract:Safe reinforcement learning (Safe RL) learns robotic controllers that optimize task rewards under safety constraints, yet observation perturbations can induce safety violations. Existing safety-directed attacks often require access to victim networks, gradients, critics, or explicit specifications -- assumptions rarely met once a controller is deployed as a black box. We propose a demonstration-guided observation attack for analyzing unknown Safe RL controllers. The framework recovers a state constraint and a surrogate policy through inverse constrained reinforcement learning, and learns dynamics from demonstration transitions. Their composed gradient generates bounded observation perturbations without victim parameters, gradients, or queries; demonstrations are the only victim-specific information. Across four Bullet tasks and one MetaDrive map with three budgets per victim, the attack exceeds every same-access baseline in 12 of 15 environment-budget conditions, and in 7 of those 12 it also exceeds the privileged reference attacks with access to the victim's reward and cost critics. Demonstrations released to support safe learning thus provide an attack surface for deployed black-box controllers. A defense study shows that state-adversarial regularization reduces attack cost, whereas the tested adversarial-training and demonstration-contamination schemes provide inconsistent protection.
| Comments: | 9 pages, 4 figures, 4 tables |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2602.16543 [cs.LG] |
| (or arXiv:2602.16543v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2602.16543 arXiv-issued DOI via DataCite |
Submission history
From: Jialiang Fan [view email]
[v1]
Wed, 18 Feb 2026 15:43:36 UTC (1,279 KB)
[v2]
Fri, 2 Oct 2026 16:38:23 UTC (2,495 KB)
来源:arXiv:cs.LG · arxiv.org