arXiv:cs.LG(机器学习,全量分类)· Morgan Byrd, Maks Sorokin, Robert Wright, Sehoon Ha·· 14 小时前AI 评分36
奖励即观测:学习基于奖励的策略以实现快速适应
Reward as Observation: Learning Reward-Based Policies for Rapid Adaptation
AI 导读
研究者提出一种仅以奖励和动作为条件的奖励策略,可在观测空间完全不同的源环境与目标环境之间实现零样本迁移。该策略在 Pointmass、Cartpole 和 2D Car Racing 三个环境中训练后,能零样本迁移到不同配色、3D 渲染等全新观测,以及 Habitat-Sim 中的 Stretch 机器人导航。基于奖励的策略还可进一步引导目标环境中基于观测的策略训练。
正文
Abstract:This paper explores a reward-based policy to achieve zero-shot transfer between source and target environments with completely different observation spaces. While humans can demonstrate impressive adaptation capabilities, deep neural network policies often struggle to adapt to a new environment and require a considerable amount of samples for successful transfer. Instead, we propose a novel reward-based policy only conditioned on rewards and actions, enabling zero-shot adaptation to new environments with completely different observations. We discuss the challenges and feasibility of a reward-based policy and then propose a practical algorithm for training. We demonstrate that a reward policy can be trained within three different environments, Pointmass, Cartpole, and 2D Car Racing, and transferred to completely different observations, such as different color palettes or 3D rendering, or Stretch robot navigation in Habitat-Sim, in a zero-shot manner. We also demonstrate that a reward-based policy can further guide the training of an observation-based policy in the target environment.
| Comments: | Website: this https URL |
| Subjects: | Machine Learning (cs.LG); Robotics (cs.RO) |
| Cite as: | arXiv:2610.00729 [cs.LG] |
| (or arXiv:2610.00729v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00729 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Morgan Byrd [view email]
[v1]
Wed, 30 Sep 2026 21:18:55 UTC (2,894 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org