arXiv:cs.LG· Yudong Lin, Haoyuan Deng, Zhuoxuan Yuan, Zaijia Yang, Yuanjiang Xue, Ziwei Wang·· 5 小时前AI 评分37
UniIntervene++:面向高效真实世界强化学习的自适应干预智能体
UniIntervene++: An Adaptive Intervention Agent for Efficient Real-World Reinforcement Learning
AI 导读
UniIntervene++ 是一个自适应干预智能体,在在线强化学习中学习在自主执行与多种辅助行为之间分配控制权,联合决定何时干预、如何干预以及何时交还控制。在五项真实世界操作任务中,它平均成功率达 89.67%,比所有基线至少高 6 个百分点,同时将人工干预降至 0.77%,相对最佳基线至少减少 94.6%。代码已在 GitHub 开源。
正文
Abstract:Online reinforcement learning (RL) enables robot policies to improve through physical interaction, but the assistance they require changes as their competence evolves. Existing intervention strategies based on offline estimates or fixed decision rules can therefore become mismatched to the current policy. To address this, we propose UniIntervene++, an adaptive intervention agent that learns to allocate control between autonomous execution and heterogeneous assisted behaviors during online RL. Specifically, UniIntervene++ first formulates the evolving RL policy, trajectory correction, and a task-structured CodePolicy as Options in a unified semi-Markov decision process and learns their relative values online. Building on this, competence-adaptive intervention periodically probes the RL policy through unassisted execution, keeping control allocation responsive to its evolving capability. Finally, coupled experience learning allows assisted behaviors to improve the RL policy, whose evolving outcomes in turn reshape future intervention decisions. In this way, UniIntervene++ jointly determines when to intervene, how to intervene, and when to return control as the RL policy improves. Across five real-world manipulation tasks, UniIntervene++ achieves an average success rate of 89.67%, outperforming all baselines by at least 6 percentage points, while reducing human intervention to 0.77%, a relative reduction of at least 94.6% from the best baseline. Code is available in our \href{this https URL}{GitHub repository}.
| Comments: | Yudong Lin and Haoyuan Deng contributed equally. Ziwei Wang is the corresponding author. Code is available in our \href{this https URL}{GitHub repository} |
| Subjects: | Machine Learning (cs.LG); Robotics (cs.RO) |
| Cite as: | arXiv:2610.03620 [cs.LG] |
| (or arXiv:2610.03620v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03620 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Haoyuan Deng [view email]
[v1]
Fri, 2 Oct 2026 17:18:12 UTC (19,520 KB)
来源:arXiv:cs.LG · arxiv.org