arXiv:cs.LG· Suzannah Wistreich, Stephen Tian, Isabella Huang, Vitor Campagnolo Guizilini, Sergey Zakharov, Katherine Liu, Jiajun Wu·· 4 小时前AI 评分35
MobileVISTA:面向移动操作姿态泛化的生成式数据增强框架
MobileVISTA: Generative Data Augmentation for Pose Generalization in Mobile Manipulation
AI 导读
MobileVISTA 是一个数据生成框架,通过联合增强自我中心视觉观测和重定向动作来补偿基座姿态变化,将标准姿态下的演示数据转化为多样化的姿态扰动训练数据。该框架针对相机随运动链移动的自我中心平台(如人形机器人)设计,在仿真人形与双臂任务以及真实 Galaxea R1 Pro 上验证,无需额外采集演示或训练生成模型即可提升策略对分布外姿态的鲁棒性,且在人形机器人上收益最大。
正文
Abstract:Mobile manipulators such as humanoid robots are increasingly deployed in dynamic, unstructured environments to perform dexterous manipulation tasks. However, end-to-end manipulation policies trained to imitate demonstration data collected from a single robot pose are brittle: even centimeter-scale deviations in robot pose at deployment can drive ego-centric observations and end-effector trajectories out of the training distribution, leading to sharp drops in performance. We introduce MobileVISTA, a data generation framework that transforms demonstrations captured at canonical poses into diverse, pose-perturbed training data by jointly (1) augmenting egocentric visual observations and (2) retargeting actions to compensate for base pose changes. Unlike prior methods, which assume a camera rigidly mounted off the actuated chain or non-trivial articulated robot geometry largely out of frame, MobileVISTA targets compatibility with egocentric platforms (e.g., humanoids) where the camera is both influenced by and must observe the robot's kinematic chain as it moves. We study MobileVISTA in simulated tasks spanning humanoid and bimanual embodiments, and on a real Galaxea R1 Pro. We find policies trained on MobileVISTA-augmented data demonstrate improved robustness to previously out-of-distribution poses encountered at test time, without additional demonstration collection or a trained generative model. Additionally, we find MobileVISTA's benefit is largest on tested humanoids, where the camera rides the actuated chain and the robot fills much of the frame. Additional videos and appendix can be found on our website: this https URL
| Subjects: | Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07511 [cs.RO] |
| (or arXiv:2610.07511v1 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07511 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Suzannah Wistreich [view email]
[v1]
Mon, 5 Oct 2026 23:24:26 UTC (1,402 KB)
来源:arXiv:cs.LG · arxiv.org