跳到正文
arXiv:cs.LG· Fengxu Liu, Siwei Wang, Gal Dalal, Shie Mannor, Yihan Du·· 4 小时前AI 评分26

线性函数逼近下基于分段奖励反馈的强化学习研究

Reinforcement Learning with Segment Reward Feedback under Linear Function Approximation

AI 导读

针对经典 RL 要求每个状态-动作对都有奖励、而轨迹级反馈又过于稀疏的问题,该研究在线性函数逼近下提出分段奖励反馈模型,为已知转移的等长分段设计了 $\bitssegd$ 与 $\edlinucbsegd$ 算法,并建立了近乎匹配的下界。

正文

View PDF HTML (experimental)

Abstract:Classical reinforcement learning (RL) assumes that a reward is observed for every visited state-action pair. However, in real-world applications such as autonomous driving, such fine-grained feedback can be costly or difficult to collect, whereas trajectory-level feedback may be too sparse for efficient learning. To provide a general feedback model bridging these two extremes and handle large state spaces, we study RL with segment reward feedback under linear function approximation. Our work answers how the granularity of segment feedback and the choice of segmentation influence learning. For equal-length segments with known transitions, we design algorithms $\bitssegd$ and $\edlinucbsegd$ for binary and sum feedback types, respectively. They adopt posterior sampling with planning to achieve computational efficiency and the E-optimal experimental design to attain near-optimality. Nearly matching lower bounds are established. For equal-length segments with unknown transitions, we develop a unified $\seglsvits$ framework with two instantiations for binary and sum feedback, which carefully integrates the posterior estimated reward parameters into least-squares value iteration. These results reveal a fundamental insight: under binary feedback, increasing the number of segments significantly reduces the regret through an exponential factor, while surprisingly, under sum feedback, the granularity of segments does not affect learning much. Finally, to investigate whether segmenting according to state-action features can further expedite learning, we design an algorithm $\uneqsegbitsd$ that allows arbitrary segmentations. The resulting regret bound shows that under the usual elliptical potential analysis, the influence of state-action features on the regret appears only through logarithmic factors, and equal segmentation achieves the best performance.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.08271 [cs.LG]
  (or arXiv:2610.08271v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.08271

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Fengxu Liu [view email]
[v1] Tue, 6 Oct 2026 12:45:39 UTC (337 KB)

来源:arXiv:cs.LG · arxiv.org