跳到正文
arXiv:cs.LG· Junpeng Yue, Boyuan Li, Yuxuan Wang, Zepeng Wang, Yuhui Fu, Feiyang Xie, Yu Zhang, Jing Zhang, Xianqi Zhang, Weibo Li, Xiaofei Zheng, Yuming Fang, Jiangxing Wang, Zongqing Lu·· 3 小时前

Being-M0.7:面向人形机器人的潜在世界-动作模型

Being-M0.7: A Latent World-Action Model for Humanoid Robots

AI 导读

Being-M0.7 是一个潜在世界-动作模型,通过预训练、机器人中训练和动作后训练三阶段,将混合模态人类数据中学到的视觉-运动先验迁移到人形机器人控制。团队构建了超 10,000 小时以人为中心的原始数据语料,融合纯视频、纯运动与配对视频-运动流。该模型在 SIMPLE 上取得对比基线中最高的综合成功率,并在真实 Unitree G1 全身移动操作任务上追平最强基线。

正文

Authors:Junpeng Yue, Boyuan Li, Yuxuan Wang, Zepeng Wang, Yuhui Fu, Feiyang Xie, Yu Zhang, Jing Zhang, Xianqi Zhang, Weibo Li, Xiaofei Zheng, Yuming Fang, Jiangxing Wang, Zongqing Lu

View PDF HTML (experimental)

Abstract:Humanoid loco-manipulation requires coordinated locomotion and manipulation informed by future scene evolution and whole-body motion, yet learning these capabilities is constrained by scarce robot demonstrations. Human video and motion datasets offer scalable supervision, but many contain only video or motion rather than paired video-motion data. Moreover, human motion does not directly specify executable robot actions. We present Being-M0.7, a latent world-action model that transfers visual-motion priors learned from mixed-modality human data to humanoid control through pre-training, robot mid-training, and action post-training. We curate a corpus from more than 10,000 hours of raw human-centric data, integrating video-only, motion-only, and paired video-motion streams to learn complementary visual dynamics and whole-body kinematic structure. Joint prediction of future latent visual states and motion encourages visual representations to encode future kinematics. Robot mid-training adapts this coarse-grained prior to robot viewpoints and body dynamics. During action post-training, an action expert combines visual predictive representations from the frozen, adapted prior with current images and proprioception through gated cross-attention, grounding predictive context in executable whole-body commands. Being-M0.7 achieves the highest aggregate success rate among the compared baselines on SIMPLE and matches the strongest baseline on real-world Unitree G1 loco-manipulation tasks.
Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as: arXiv:2610.11283 [cs.RO]
  (or arXiv:2610.11283v1 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2610.11283

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Junpeng Yue [view email]
[v1] Thu, 8 Oct 2026 05:49:24 UTC (14,294 KB)

来源:arXiv:cs.LG · arxiv.org