跳到正文
arXiv:cs.LG· Shripad V. Deshmukh, Yaswanth Chittepu, Dhawal Gupta, Philip Thomas, Scott Niekum·· 4 小时前AI 评分45

凸凹强化学习(CCRL):把策略优化写成 DC 规划并恢复 CPI、NPG、TRPO、AWR

Convex-Concave Reinforcement Learning

AI 导读

研究者提出凸凹强化学习(CCRL),在 log 密度比坐标 y := log[π/π_n] 下,用 PDIS 计算的精确单步目标属于 DC 约束的 DC 规划,可用序列凸规划(SCP)求解并给出收敛保证,CPI、NPG、TRPO、AWR 均为其特例。

正文

View PDF HTML (experimental)

Abstract:Policy learning drives many of the most consequential and heavily-invested applications of reinforcement learning today. Yet the core optimization problem it rests on (maximizing expected return) is notoriously non-convex, even under a direct policy parameterization, and the field has largely responded by avoiding it: optimizing convex surrogate approximations of the return under trust-region constraints (NPG, TRPO, PPO, AWR). We show that this seemingly unstructured problem is not actually structureless. In log-density-ratio coordinates $y := \log[\pi/\pi_n]$, the exact per-iteration objective, computable via per-decision importance sampling (PDIS), is a difference-of-convex-constrained difference-of-convex (DC-constrained DC) program. This structure lets us move beyond surrogate approximations: it recovers CPI, NPG, TRPO, and AWR as special cases along interpretable axes, and it opens a multi-step axis $k$ that couples consecutive decisions. We solve the per-iteration program with sequential convex programming (SCP), the standard solver for difference-of-convex problems, and give convergence guarantees under mild conditions, bridging the difference-of-convex optimization and RL literatures. Empirically, multi-step Convex-Concave RL (CCRL) wins on diagnostic MDPs where credit must propagate across a horizon (its advantage growing with the dependency length), is competitive with a tuned PPO on classic control, and on a realistic, stochastic, mid-horizon healthcare domain converges markedly faster than tuned PPO to the same near-optimal survival, with an 11.3% higher area under the training curve.
Comments: 37 pages, 5 figures. Code: this https URL
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.09108 [cs.LG]
  (or arXiv:2610.09108v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09108

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Shripad Deshmukh [view email]
[v1] Tue, 6 Oct 2026 20:59:21 UTC (280 KB)

来源:arXiv:cs.LG · arxiv.org