跳到正文
arXiv:cs.LG· Huizhen Yu, Isaiah Heidt·· 4 小时前AI 评分34

面向多链 MDP 的平均奖励强化学习:一种分层分解方法

Average-Reward Reinforcement Learning for Multichain MDPs: A Hierarchical Decomposition Approach

AI 导读

研究团队提出基于异步值迭代的强化学习算法,用于求解平均奖励多链 MDP 的最优策略,仅需知道 MDP 转移图,并借助 Bather 分解将状态空间分层划分为通信子系统与瞬态状态。

正文

View PDF HTML (experimental)

Abstract:We study learning optimal policies in average-reward multichain Markov decision processes (MDPs), where the optimal gain may depend on the initial state and recurrence structures vary across policies, creating challenges for reinforcement learning (RL) methods. We propose an asynchronous value-iteration-based RL algorithm that requires no model knowledge beyond the MDP's transition graph and leverages Bather's decomposition to hierarchically partition the state space into communicating subsystems and transient states. This decomposition induces a recasting of the global decision problem into structured subproblems, which our algorithm exploits. We show that the algorithm converges to the optimal gain and produces gain-optimal policies after finite time. Building on this base algorithm, we develop two further algorithms: one approximately solves the multichain average optimality equations to obtain near gain-optimal policies, and another targets near bias-optimality by approximating the optimal bias function and solving an induced average-reward multichain MDP using the base algorithm. We provide almost-sure convergence guarantees for all three algorithms and empirically compare their tradeoffs, showing that the latter two also consistently improve transient performance relative to the base algorithm. To our knowledge, these are the first essentially model-free average-reward RL algorithms for general multichain MDPs without reductions to discounted problems.
Comments: 60 pages, 4 figures
Subjects: Machine Learning (cs.LG); Optimization and Control (math.OC)
MSC classes: 90C40, 68T05, 93E20
Cite as: arXiv:2610.10326 [cs.LG]
  (or arXiv:2610.10326v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.10326

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Huizhen Yu [view email]
[v1] Wed, 7 Oct 2026 16:13:15 UTC (734 KB)

来源:arXiv:cs.LG · arxiv.org