arXiv:cs.LG· Guhyeon Kang, Minhae Kwon·· 4 小时前AI 评分33
VAN-Flow:面向稀疏长时程环境的方差规避 n 步离线强化学习
Variance-Averse $n$-Step Offline Reinforcement Learning for Sparse Long-Horizon Environments
AI 导读
VAN-Flow 是一个面向生成式离线强化学习的框架,通过类别分布式评论家、方差规避期望算子和流匹配生成演员三部分,在 D4RL 和 OGBench 的 40 多个任务上持续超越强基线,长时程与高方差场景增益最大。该期望算子对类别回报分布的概率质量进行平滑重加权,偏好高回报且低离散度的动作,无需硬截断或辅助惩罚项。论文已被 NeurIPS 2026 接收。
正文
Abstract:Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions. However, this expressiveness also exposes a key challenge in heterogeneous datasets: generative policies can reproduce unreliable action modes whose return distributions exhibit high variance, occasionally yielding high returns by chance but lacking consistency. Consequently, maximizing the expected $Q$-value alone is insufficient for identifying reliable actions. We propose VAN-Flow (Variance-Averse $n$-step Flow), a framework that promotes reliable actions in generative offline RL. VAN-Flow combines (i) a categorical distributional critic, (ii) a variance-averse expectation operator that smoothly reweights atom probabilities to favor actions with both high returns and low dispersion, and (iii) a flow-matching generative actor guided via rejection sampling. Unlike CVaR or mean-variance objectives, the operator redistributes probability mass over the categorical return distribution without hard truncation or auxiliary penalty terms. Across more than 40 tasks from D4RL and OGBench, VAN-Flow consistently outperforms strong baselines, with the largest gains in long-horizon and high-variance regimes where reliable action selection becomes critical.
| Comments: | Accepted at NeurIPS 2026 |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07899 [cs.LG] |
| (or arXiv:2610.07899v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07899 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Guhyeon Kang [view email]
[v1]
Tue, 6 Oct 2026 07:47:26 UTC (1,480 KB)
来源:arXiv:cs.LG · arxiv.org