arXiv:cs.LG(机器学习,全量分类)· Sichang Su, Benjamin Yang, Zhiyun Deng, Boyuan Liang, Yip Fun Yeung, Zelin Wang, Lingfeng Sun·· 1 天前AI 评分43
PARTS:面向长时程机器人操作的实世界子任务 RL 框架
From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention
AI 导读
PARTS 是一个实世界子任务强化学习框架,通过冻结预训练策略提供基础动作、由智能体生成选择器与成功验证器激活残差修正并给出局部结果奖励,让训练在极少人工干预下聚焦瓶颈子任务。
正文
Abstract:A pretrained robot foundation policy may execute most of a long-horizon task yet repeatedly fail at a few critical subtasks. Collecting additional full-task demonstrations for supervised fine-tuning (SFT) requires operators to repeat behaviors the policy already performs well. Reinforcement learning (RL) fine-tuning offers a promising path to bridge this gap, but existing approaches struggle to solve long-horizon tasks using only sparse rewards. We present PARTS (Policy Adaptation with RL on Targeted Subtasks), a real-world subtask RL framework that concentrates practice at these bottlenecks while allowing training rollouts to proceed with minimal human intervention. The frozen pretrained policy supplies nominal actions throughout execution, while agent-generated selectors and success verifiers activate residual corrections and provide local outcome rewards. These rewards support learning from successful subtasks even when complete-task successes are scarce. Training combines online RL with success-reweighted retraining, and each retrained residual policy is redeployed to collect further experience. Humans identify bottlenecks during setup and perform physical resets when needed. On bimanual YAM and single-arm Franka tasks, PARTS improves complete-task success from 32% to 61% and from 50% to 95%, respectively, using tens of minutes of real-world RL rollouts per task on average. Compared with existing real-world RL fine-tuning methods, PARTS raises full-task success by more than 25% under the same robot-rollout budget while requiring less human involvement.
| Comments: | Project page: this https URL |
| Subjects: | Robotics (cs.RO); Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.21788 [cs.RO] |
| (or arXiv:2609.21788v2 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2609.21788 arXiv-issued DOI via DataCite |
Submission history
From: Sichang Su [view email]
[v1]
Fri, 18 Sep 2026 14:03:49 UTC (1,565 KB)
[v2]
Wed, 30 Sep 2026 21:08:34 UTC (1,919 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org