跳到正文
arXiv:cs.LG· Nishanth Arun Rao, Royina Karegoudra Jayanth, Benjamin Eysenbach, Jaime Fern\'andez Fisac·· 3 小时前

面向安全关键强化学习的统一 Bellman 算子

A Unified Bellman Operator for Safety-Critical Reinforcement Learning

AI 导读

研究者提出一种统一性能与安全目标的 Bellman 算子,将二者整合进联合价值函数,并在双时间尺度随机逼近框架下证明了时序差分学习的收敛性。该框架在快时间尺度估计学习策略的安全价值、慢时间尺度估计联合价值,收敛后的最优策略在最大化任务回报的同时始终保持安全。连续控制任务的神经逼近实验显示其收敛稳定,测试时安全违规接近零。

正文

View PDF HTML (experimental)

Abstract:Reinforcement learning in safety-critical domains requires maximizing task performance while strictly adhering to safety constraints. Existing safe reinforcement learning paradigms typically force a trade-off: they either require a priori knowledge to provide strict safety guarantees (e.g., safety filters), or they enable joint learning but only satisfy safety constraints on average. In this work, we propose a novel Bellman operator that unifies performance and safety objectives into a joint value function. We show that temporal difference learning with the joint Bellman operator converges under a two-timescale stochastic approximation framework. On the fast timescale, the safety value of the learning joint policy is estimated, while the joint value is estimated on the slow timescale. Convergence is ensured by formulating the limiting dynamics as an occupation-averaged differential inclusion, and showing that it asymptotically converges to a set of limiting optimal safety-constrained task value functions. Theoretically, once converged, the resulting optimal policy maximizes task return while maintaining safety at all times. Empirical evaluations on continuous control tasks with neural approximations demonstrate stable convergence with near-zero safety violations at test time.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.12420 [cs.LG]
  (or arXiv:2610.12420v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.12420

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Nishanth Arun Rao [view email]
[v1] Thu, 8 Oct 2026 17:53:02 UTC (3,527 KB)

来源:arXiv:cs.LG · arxiv.org