跳到正文
arXiv:cs.LG· Qi Feng, Gu Wang·· 5 小时前AI 评分31

面向松弛随机控制问题的连续策略与价值迭代算法及其收敛性

Continuous Policy and Value Iteration for Relaxed Stochastic Control Problems and Its Convergence

AI 导读

研究者提出一种连续策略-价值迭代算法,通过 Langevin 型动力学同时更新随机控制问题的价值函数近似与最优控制。该工作针对无限时域、熵正则化的松弛控制问题,在哈密顿量单调性条件下证明了策略改进与收敛到最优控制。数值实验涵盖非凹奖励的 LQ 模型和一个最优控制服从偏斜拉普拉斯分布的非 LQ 示例。

正文

View PDF HTML (experimental)

Abstract:We introduce a continuous policy-value iteration algorithm where the approximations of the value function of a stochastic control problem and the optimal control are simultaneously updated through Langevin-type dynamics. This paper focuses on the entropy-regularized relaxed control problem with infinite horizon. We establish policy improvement and demonstrate convergence to the optimal control under the monotonicity condition of the Hamiltonian. By utilizing Langevin-type stochastic differential equations for continuous updates along the policy iteration direction, our approach enables the use of distribution sampling techniques in machine learning to optimize the value function and identify the optimal control simultaneously. Numerical experiments are presented for LQ model with a non-concave reward, and a non-LQ example with a skewed-Laplace optimal control.
Comments: 34 pages, 2 figures. Revised to focus exclusively on the relaxed control problem
Subjects: Optimization and Control (math.OC); Machine Learning (cs.LG)
MSC classes: 93E20, 93E35, 60H10
Cite as: arXiv:2506.08121 [math.OC]
  (or arXiv:2506.08121v3 [math.OC] for this version)
  https://doi.org/10.48550/arXiv.2506.08121

arXiv-issued DOI via DataCite

Submission history

From: Qi Feng [view email]
[v1] Mon, 9 Jun 2025 18:20:21 UTC (30 KB)
[v2] Mon, 13 Jul 2026 21:00:23 UTC (43 KB)
[v3] Thu, 1 Oct 2026 19:41:28 UTC (1,449 KB)

来源:arXiv:cs.LG · arxiv.org