跳到正文
arXiv:cs.AI· Chenzhi Liu, Yue Zhang, Jiehong Lin, Jianan Wang, Bo Wang, Zhongrui Wang, Xiaojuan Qi·· 4 小时前AI 评分48

MobiAgent:面向长时程移动操作的双环递归策略自改进框架

MobiAgent: Dual-Loop Recursive Policy Self-Improvement for Long-Horizon Mobile Manipulation

AI 导读

MobiAgent 是一个双环智能体框架,通过内环的可组合原子技能与 VLM 重规划实现鲁棒部署执行,外环则自动分割、验证并聚类部署轨迹以持续微调技能库,无需人工标注。

正文

View PDF HTML (experimental)

Abstract:Long-horizon mobile manipulation presents significant challenges due to compounding execution errors and capacity interference between locomotion and arm control. While recent Vision-Language-Action models excel at short-horizon tasks, they lack the hierarchical reasoning required for multi-stage objectives. Furthermore, existing hierarchical agents suffer from rigid sub-task mapping, inflexible replanning, and a lack of continuous learning. To address these limitations, we introduce MobiAgent, a dual-loop agentic framework that bridges robust deployment execution and recursive policy self-improvement. During deployment, the Inner Loop decouples high-level reasoning from low-level control through highly composable atomic skills. It employs Vision-Language models for receding-horizon planning and visual reflection, dynamically composing skills to ensure robust error recovery. These skills are executed by specialized flow-matching experts that share a unified VLM backbone, maximizing reusability while mitigating capacity interference. Concurrently, the Outer Loop drives automated lifelong learning by autonomously segmenting and verifying deployment rollouts, clustering them to discover atomic skills, and continuously fine-tuning the skill library without human annotations. Evaluations on RoboCasa, BEHAVIOR-1K, and real-world tasks demonstrate the effectiveness of MobiAgent. It outperforms $\pi_{0.5}$-TA by 22.5 percentage points on BEHAVIOR-1K and enables robust recovery from execution failures. Through autonomous data recycling, success improves from 7.50% to 27.50% on RoboCasa and from 32.5% to 57.5% on Astribot S1.
Comments: Accepted at the Conference on Robot Learning (CoRL) 2026. Project page: this https URL
Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.03476 [cs.RO]
  (or arXiv:2610.03476v1 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2610.03476

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Chenzhi Liu [view email]
[v1] Fri, 2 Oct 2026 15:48:23 UTC (8,608 KB)

来源:arXiv:cs.AI · arxiv.org