arXiv:cs.LG· Shengfan Cao, Francesco Borrelli·· 4 小时前AI 评分37
交错投影梯度下降:安全模仿学习的新方法
Interleaved Projected Gradient Descent for Safe Imitation Learning
AI 导读
研究者提出一种面向状态与输入约束下神经网络控制策略的模仿学习方案,训练时在标准模仿梯度步之间交替进行 k 步安全步骤,将网络动作拉向其安全集投影,运行时仅用训练好的网络、无需安全滤波器。在非线性自动驾驶赛车任务中,该方法将违规回合占比从 15% 降至 1%,约为最优权重惩罚方法违规率的六分之一,圈速与无约束模仿相当,但代价是额外训练计算量。
正文
Abstract:We propose an imitation-learning design for neural-network control policies under state and input constraints. Training alternates a standard imitation gradient step with a block of $k$ safety steps that pull the network's actions toward their projection onto the safe set; at run time, the controller is the trained network alone, with no safety filter. We analyze this scheme as inexact projected gradient descent in the space of policy actions. When the projected actions are recomputed at every safety step and each step moves the actions consistently toward the safe set, letting $k$ grow logarithmically yields asymptotic constraint satisfaction on the training states and bounds the distance to the constrained optimum of the imitation loss; with the projected actions held fixed, the same holds only if they are exactly representable by the network. On a nonlinear autonomous racing task, we compare our method with adding a weighted constraint-violation penalty to the imitation loss. With a sufficiently large weight, our method matches the lap time of unconstrained imitation while reducing the fraction of violating episodes from $15\%$ to $1\%$, about six times fewer than the penalty approach at its best weight. Its lap times are less sensitive to the weight, which instead sets how quickly violations vanish during training. In racing, the safety corrections are sparse and the conditions of the analysis do not hold; the gain arises instead through the data collected during training. These gains come at the cost of additional training computation.
| Comments: | Submitted to the 2027 American Control Conference (ACC). 9 pages, 3 figures |
| Subjects: | Systems and Control (eess.SY); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07167 [eess.SY] |
| (or arXiv:2610.07167v1 [eess.SY] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07167 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shengfan Cao [view email]
[v1]
Mon, 5 Oct 2026 18:00:08 UTC (593 KB)
来源:arXiv:cs.LG · arxiv.org