arXiv:cs.LG· Mahmoud Selim, Cristina Cipriani, Karl Henrik Johansson·· 4 小时前AI 评分37
ICDP:面向自动驾驶的交互约束离线强化学习框架
Beyond Policy Support: Interaction Constrained Offline Reinforcement Learning for Autonomous Driving
AI 导读
针对自动驾驶离线强化学习中"交互分布偏移"(IDS)导致价值估计不可靠的问题,研究者提出交互约束驾驶策略ICDP框架,通过对比密度比估计分离出交互兼容性,无需联合密度建模或仿真器推演即可控制交互级分布偏移。在nuPlan、InterPlan及真实卡车实验的闭环评估中,ICDP抑制了高价值但交互不支持轨迹的选择,并提升了交互关键场景下的驾驶性能。
正文
Abstract:Offline reinforcement learning enables reward-driven policy improvement from fixed datasets without requiring online exploration, making it particularly attractive in safety-critical domains. A central challenge, however, is distribution shift: policy optimization may favor actions that are weakly supported by the offline data, rendering value estimates unreliable. Existing approaches primarily control this shift in the policy's own action space. In interactive environments such as autonomous driving, this can be insufficient: a candidate ego trajectory may remain well supported under the marginal behavior distribution while being poorly supported jointly with the surrounding-agent behavior observed in the logged interaction. We refer to this degradation in interaction support as \emph{interaction distribution shift} (IDS), and introduce \emph{Interaction-Constrained Drive Policy} (ICDP), an offline reinforcement learning framework that explicitly controls interaction-level distribution shift. Starting from the joint data distribution over ego and surrounding-agent futures, we show that joint-support degradation decomposes exactly into an ego-support component and a residual interaction-support component. We recover the latter through contrastive density-ratio estimation, isolating interaction compatibility without explicit joint-density modeling, surrounding-agent prediction, or rollouts in reactive simulators or learned world models during policy optimization. Closed-loop evaluations on nuPlan, Interplan and real-world truck experiments show that ICDP suppresses high-value yet interaction-unsupported trajectory selections and improves performance in interaction-critical driving scenarios. Project webpage: this https URL
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO) |
| Cite as: | arXiv:2610.09763 [cs.LG] |
| (or arXiv:2610.09763v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09763 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mahmoud Selim [view email]
[v1]
Wed, 7 Oct 2026 09:48:53 UTC (4,687 KB)
来源:arXiv:cs.LG · arxiv.org