arXiv:cs.LG· Shangzhe Li, Weitong Zhang·· 3 小时前
从离线到在线强化学习的价值自适应:O2O-LSVI 实现可证明的高效微调
Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation
AI 导读
研究者提出 O2O-LSVI 算法,在通用函数逼近下解决离线到在线强化学习的价值自适应问题,并证明在特定结构条件下其样本复杂度优于纯在线 RL。研究还建立了 minimax 下界,表明当预训练 Q 函数接近最优时,在线自适应在某些困难实例上仍无法比纯在线 RL 更高效。神经网络实验验证了该方法的实际有效性。
正文
Abstract:We study value adaptation in offline-to-online reinforcement learning under general function approximation. Starting from an imperfect offline pretrained $Q$-function, the learner aims to adapt it to the target environment using only a limited amount of online interaction. We first characterize the difficulty of this setting by establishing a minimax lower bound, showing that even when the pretrained $Q$-function is close to optimal $Q^\star$, online adaptation can be no more efficient than pure online RL on certain hard instances. On the positive side, under a novel structural condition on the offline-pretrained value functions, we propose O2O-LSVI, an adaptation algorithm with problem-dependent sample complexity that provably improves over pure online RL. Finally, we complement our theory with neural-network experiments that demonstrate the practical effectiveness of the proposed method.
| Comments: | 44 pages, 2 tables |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2604.13966 [cs.LG] |
| (or arXiv:2604.13966v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2604.13966 arXiv-issued DOI via DataCite |
Submission history
From: Shangzhe Li [view email]
[v1]
Wed, 15 Apr 2026 15:17:40 UTC (37 KB)
[v2]
Wed, 7 Oct 2026 22:02:17 UTC (87 KB)
来源:arXiv:cs.LG · arxiv.org