跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Jia Wan, Sean R. Sinclair, Devavrat Shah, Martin J. Wainwright·· 15 小时前AI 评分31

Exo-MDPs 如何实现样本高效的强化学习:结构等价与极小极大遗憾分析

Exploiting Exogenous Structure for Sample-Efficient Reinforcement Learning

AI 导读

研究针对 Exo-MDPs 这类状态空间分为外生与内生组件的马尔可夫决策过程,建立了离散 MDP、Exo-MDPs 与离散线性混合 MDP 之间的表示等价关系。

正文

View PDF HTML (experimental)

Abstract:We study a structured class of Markov Decision Processes, known as Exo-MDPs, in which the state space is partitioned into exogenous and endogenous components. Exogenous states evolve stochastically, independent of the agent's actions, while endogenous states evolve deterministically based on both state components and actions. Exo-MDPs capture many operations research settings, including inventory control, resource management, and ride-sharing. Our first contribution is structural: we establish a representational equivalence between discrete MDPs, Exo-MDPs, and discrete linear mixture MDPs. Our second contribution is statistical. We characterize the minimax regret of learning in Exo-MDPs when the effective dimension r is small relative to the endogenous state and action spaces. When the exogenous states are unobserved, we prove matching upper and lower regret bounds of order $\Theta(Hr \sqrt{K})$ over $K$ episodes of horizon $H$, where $r$ is the effective dimension of the Exo-MDP. When exogenous states are observed, the minimax regret improves to $\Theta(H\sqrt{ r K})$, revealing a $\Theta(\sqrt{r})$ statistical gap due to observation of the exogenous states. These results show that Exo-MDPs decouple sample complexity from action space and endogenous state space. We validate these insights with experiments on inventory control and resource allocation.
Comments: 76 pages
Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG); Optimization and Control (math.OC)
Cite as: arXiv:2409.14557 [stat.ML]
  (or arXiv:2409.14557v5 [stat.ML] for this version)
  https://doi.org/10.48550/arXiv.2409.14557

arXiv-issued DOI via DataCite

Submission history

From: Sean R. Sinclair [view email]
[v1] Sun, 22 Sep 2024 18:45:38 UTC (117 KB)
[v2] Mon, 14 Oct 2024 23:46:09 UTC (152 KB)
[v3] Wed, 5 Feb 2025 15:49:21 UTC (768 KB)
[v4] Thu, 23 Jul 2026 13:51:31 UTC (251 KB)
[v5] Wed, 30 Sep 2026 20:32:12 UTC (209 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org