跳到正文
arXiv:cs.LG· Zihao Zhao, Ashwath K. Karunakaram, Ali Eshragh, Yuexing Li, Kai Wang·· 4 小时前AI 评分36

MDP 中的决策聚焦学习:一种占用度量方法

Decision-Focused Learning in MDPs: An Occupancy Measure Approach

AI 导读

研究者提出用占用度量(occupancy measure)将 MDP 重构为线性规划(LP),通过 pivoting 算法识别有效约束得到闭式梯度,替代需在全状态-动作对上求解线性系统的 KKT 方法。

正文

View PDF HTML (experimental)

Abstract:In this work, we consider decision-focused learning (DFL) for a Markov decision process (MDP), where existing methods differentiate through the KKT conditions of the Bellman equation and require solving a linear system over all state-action pairs, limiting its scalability. We address this by reformulating the MDP as an occupancy measure-based linear program (LP), whose feasible region is induced by predicted dynamics, and we derive a closed-form gradient by identifying the active constraints in the feasible polyhedron via the pivoting algorithm. This occupancy measure-based LP layer raises two challenges: (1) LP's solution gradient is discontinuous when active constraints change, and (2) the LP backward cost still scales with the state size, which is costly for large or continuous state spaces. We address the challenges with an augmented Lagrangian surrogate and smooth the boundary jumps by random row sketching of the constraints, and a learnable soft state-aggregation layer and its function-approximation generalization that scales the LP to large finite and continuous-state MDPs. Across multiple tasks, our methods reach lower regret than KKT-based DFL and two-stage baselines with significantly lower computation cost. The source code for all experiments is available at this https URL.
Comments: Accepted at NeurIPS 2026
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.08384 [cs.LG]
  (or arXiv:2610.08384v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.08384

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zihao Zhao [view email]
[v1] Tue, 6 Oct 2026 14:04:19 UTC (535 KB)

来源:arXiv:cs.LG · arxiv.org