arXiv:cs.LG· Iris Xu, Sunshine Jiang, John Marangola, Pulkit Agrawal, Zhang-Wei Hong·· 7 小时前AI 评分46
Learning from Hindsight:让 VLA 强化学习从失败中学习
Learning from Hindsight for VLA Reinforcement Learning
AI 导读
研究者提出 Learning from Hindsight(LfH),用预训练视觉语言模型为失败的 rollout 重新标注其实际完成的行为,并将这些辅助任务与指令任务联合训练 VLA 策略。
正文
Abstract:Reinforcement learning is increasingly used to fine-tune vision-language-action (VLA) models, but robot interaction is expensive and learning becomes highly sample inefficient when successful rollouts are rare. When reward is assigned only for completing the commanded task, a failed rollout is treated as having no value even if it successfully executes behaviors relevant to that task. A robot that fails to place the correct object in a bowl may still move that object toward the bowl or place a different object inside it, demonstrating objects and actions that can be reused to solve the target task. These behaviors define auxiliary tasks that the policy can already solve, providing useful learning signals even before it can solve the harder target task. We introduce $\textit{Learning from Hindsight (LfH)}$, which turns such failures into additional learning signals. Using a pretrained vision-language model, LfH relabels failed rollouts with the behaviors they actually accomplish and trains the policy jointly on the commanded task and these auxiliary tasks. On out-of-distribution LIBERO-PRO manipulation tasks, LfH matches the final performance of GRPO with approximately $5\times$ fewer rollouts and improves sample efficiency across multiple VLA backbones. On a physical Franka robot, LfH raises success from $0\%$ to $56\%$ within 160 training rollouts, while GRPO reaches $22\%$.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.09042 [cs.LG] |
| (or arXiv:2607.09042v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2607.09042 arXiv-issued DOI via DataCite |
Submission history
From: Iris Xu [view email]
[v1]
Fri, 10 Jul 2026 02:17:41 UTC (2,010 KB)
[v2]
Mon, 5 Oct 2026 20:49:25 UTC (1,851 KB)
来源:arXiv:cs.LG · arxiv.org