跳到正文
arXiv:cs.LG· Valliappan Chidambaram Adaikkappan, Sai Rajeswar, Pietro Mazzaglia, David Meger·· 7 小时前AI 评分32

MSPR:面向目标条件强化学习的多尺度预测表示

MSPR: Multi-scale Predictive Representations for Goal-conditioned Reinforcement Learning

AI 导读

MSPR 是一个面向离线目标条件强化学习(GCRL)的多尺度预测监督框架,通过捕获从局部物理动态到长程目标结构的对齐目标,约束隐空间中状态与目标的对齐,缓解稀疏奖励下编码器学到目标无关特征的问题。MSPR 在视觉与状态类任务上均取得强劲表现,并在具有挑战性的数据条件下保持 SOTA 性能。

正文

View PDF HTML (experimental)

Abstract:This paper investigates robust representation learning in offline goal-conditioned reinforcement learning (GCRL). Particularly in sparse reward scenarios, learning representations that align state and goal latents is a challenge, as the encoder can learn goal-agnostic features that destabilize policy learning. We address this issue by learning the encoder's representation with alignment objectives that capture the environment across multiple scales, from local physical dynamics to long-horizon goal-directed structure. Concretely, we propose MSPR, a framework that leverages multi-scale predictive supervision to enforce goal-directed alignment within the latent space. We demonstrate that MSPR leads to strong performance on both vision and state-based tasks. Furthermore, we show that our approach is resilient under realistic, challenging data regimes, maintaining state-of-the-art performance across a wide variety of tasks.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2605.09364 [cs.LG]
  (or arXiv:2605.09364v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2605.09364

arXiv-issued DOI via DataCite

Submission history

From: Valliappan Chidambaram Adaikkappan [view email]
[v1] Sun, 10 May 2026 06:27:20 UTC (2,737 KB)
[v2] Tue, 6 Oct 2026 17:05:14 UTC (4,184 KB)

来源:arXiv:cs.LG · arxiv.org