跳到正文
arXiv:cs.AI· Bohan Zhou, Xingbei Chen, Emily Huang, Weilin Ruan, Haojian Huang, Yehang Zhang, Zexi Li, Wenqian Li, Qize Yu, Zetian Song, Leyi Wu, Jinghao Li, Mingxuan Song, Xinrun Xu, Zongyang Qiu, Yangkai Wei, Tianyi Zhang, Kaiwen Zhou, Yinchuan Li, James Cheng·· 3 小时前

RoboAware:从反事实结果中学习协调具身技能

RoboAware: Learning to Coordinate Embodied Skills from Counterfactual Outcomes

AI 导读

RoboAware 通过仅学习状态条件化的责任协调器,从反事实结果中协调模块化机器人技能与冻结的端到端策略,在 100 项任务上达到 77.0% 的整体成功率,RoboSuite 上平均 90.0%、LIBERO-Pro 任务簇 73.8%、RoboTwin 双臂任务 90.0%,超越现有 code-as-policy 与 VLA-harness 基线。

正文

Authors:Bohan Zhou, Xingbei Chen, Emily Huang, Weilin Ruan, Haojian Huang, Yehang Zhang, Zexi Li, Wenqian Li, Qize Yu, Zetian Song, Leyi Wu, Jinghao Li, Mingxuan Song, Xinrun Xu, Zongyang Qiu, Yangkai Wei, Tianyi Zhang, Kaiwen Zhou, Yinchuan Li, James Cheng

View PDF HTML (experimental)

Abstract:Embodied coding agents can combine modular robot skills with frozen end-to-end policies, yet effective composition requires anticipating which policy family will succeed in the current physical state. We present RoboAware, which builds on coding agents' skill orchestration by learning only a state-conditioned responsibility coordinator from counterfactual outcomes. Inspired by the success of REPL, we propose the $P^5$ schema and formulate a hierarchical MDP based on it. $P^5$ organizes skills uniformly into five semantic stages, defining where responsibility can be compared. To address the lack of counterfactual branch outcomes in existing work, we introduce State-Locked Counterfactual Branching (SCB), which restores the same training state to generate and execute a code block from each admissible family, exposing outcomes that selected-branch experience leaves unobserved. Building on this, we propose Execution-Aware Learning (EAL), which combines Monte Carlo tree search with Q-learning to distill these outcomes into family-conditioned values. At deployment, the coordinator selects the policy family according to observable context, and the frozen coding agent generates the next local code block. Comprehensive single-episode evaluations on 100 tasks show that RoboAware reaches a 77.0% overall success rate, with SOTA averages of 90.0% on RoboSuite, 73.8% on diverse LIBERO-Pro task clusters, and 90.0% on challenging RoboTwin bimanual tasks, outperforming existing code-as-policy and VLA-harness baselines.
Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11480 [cs.RO]
  (or arXiv:2610.11480v1 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2610.11480

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Bohan Zhou [view email]
[v1] Thu, 8 Oct 2026 08:25:49 UTC (1,457 KB)

来源:arXiv:cs.AI · arxiv.org