跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Jiaxi Ye, Chunji Lv, Guoren Wang, Changsheng Li·· 15 小时前AI 评分33

RADP:面向自动驾驶的规则对齐扩散规划,让规划过程可解释

Learning to Explain While Planning: Rule-Aligned Diffusion Planning for Autonomous Driving

AI 导读

研究者提出规则对齐扩散规划器 RADP,将可微驾驶规则融入扩散训练目标,使规则知识成为超越有限专家演示的内在行为准则。同时提出 Rule-Pressure Attribution(RPA),用规则损失对预测轨迹的梯度构建监督信号,在线估计各规则的优化压力。在 nuPlan 上,RADP 提升了安全关键场景的闭环规划表现,RPA 的规则压力与后续规则特定风险保持时间对齐。

正文

View PDF HTML (experimental)

Abstract:Diffusion planners exhibit strong capabilities in generating multimodal trajectories. However, existing methods primarily rely on expert demonstrations to fit trajectory distributions, learning statistical correlations among scenes, behaviors, and trajectories without explicitly modeling driving rules. In long-tail scenarios where expert data are scarce, the lack of behaviors to imitate may lead to trajectories that violate safety or compliance requirements. Moreover, their generation process lacks rule-level explanations, making it difficult to determine which rules drive trajectory adjustments, when they take effect, and how strongly they act, thereby limiting failure diagnosis, safety validation, and targeted improvement. To address these limitations, we propose the Rule-Aligned Diffusion Planner (RADP), which incorporates differentiable driving rules into the diffusion objective during training, turning rule knowledge into intrinsic behavioral principles beyond finite demonstrations. We further introduce Rule-Pressure Attribution (RPA), which constructs supervision signals from gradients of rule losses with respect to predicted trajectories and employs a lightweight attribution head to estimate the optimization pressure exerted by each rule online. To assess the closed-loop behavioral relevance of these attributions, we propose a temporal risk-alignment protocol that evaluates whether current rule pressures reflect corresponding risks during subsequent closed-loop execution. Experiments on nuPlan show that RADP improves closed-loop planning in challenging safety-critical scenarios, while RPA exhibits consistent temporal alignment with subsequent rule-specific risks, validating both intrinsic rule learning and rule-level interpretability.
Comments: Corrected the author metadata and an author-name typo; manuscript content unchanged
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.39995 [cs.LG]
  (or arXiv:2609.39995v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.39995

arXiv-issued DOI via DataCite

Submission history

From: Jiaxi Ye [view email]
[v1] Wed, 30 Sep 2026 15:46:21 UTC (12,734 KB)
[v2] Thu, 1 Oct 2026 08:34:11 UTC (6,367 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org