跳到正文
arXiv:cs.LG· Lijun Bo, Yijie Huang, Chenhao Lu·· 4 小时前AI 评分27

基于确定性策略梯度的超立方体状态空间边界感知强化学习

Boundary-aware Reinforcement Learning for Hypercube State Spaces via Deterministic Policy Gradient

AI 导读

研究者提出连续时间确定性策略梯度框架,用于状态受超立方体上受控反射随机微分方程支配的强化学习,建立了值函数与 Neumann Bellman 方程的关联,并给出确定性策略梯度公式与鞅刻画定理。基于该理论提出连续时间深度确定性策略梯度算法,通过软惩罚或硬架构约束施加 Neumann 边界条件。储层控制实验显示,边界感知方法显著降低 Neumann 边界残差并提升学习稳定性。

正文

View PDF HTML (experimental)

Abstract:We develop a continuous-time deterministic policy gradient framework for reinforcement learning with reflected state dynamics, where the state process is governed by a controlled reflected stochastic differential equation on a hypercube. Under suitable regularity assumptions, we establish the connection between the value function and the Neumann Bellman equation, introduce an advantage-rate function that yields a deterministic policy gradient formula, and prove the martingale characterization theorem. Motivated by these theoretical results, we propose a continuous-time deep deterministic policy gradient algorithm for reflected stochastic systems, in which the Neumann boundary condition is imposed via either soft penalization or hard architectural constraint. We further quantify the discrepancy between the ideal continuous-time dynamics and the discretely sampled exploratory dynamics executed in practice, showing that the error decays as the time grid is refined and exploration noise vanishes. Our experiments on reservoir control problems illustrate the effectiveness of the RL framework, highlighting that boundary-aware methods substantially reduce Neumann boundary residuals and enhance learning stability.
Subjects: Optimization and Control (math.OC); Machine Learning (cs.LG)
Cite as: arXiv:2610.09712 [math.OC]
  (or arXiv:2610.09712v1 [math.OC] for this version)
  https://doi.org/10.48550/arXiv.2610.09712

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yijie Huang [view email]
[v1] Wed, 7 Oct 2026 09:07:40 UTC (4,828 KB)

来源:arXiv:cs.LG · arxiv.org