arXiv:cs.LG· Mert Albaba, Jens Bei{\ss}wenger, Anna Manasyan, Daniel Marta, Michael J. Black, Wieland Brendel, Andreas Krause, Georg Martius, Martin Riedmiller·· 3 小时前
VioLA:从人类数据学习通用人形机器人控制策略
VioLA: Learning Generalist Humanoid Control Policies from Human Data
AI 导读
VioLA 通过预测身体与手部运动潜变量而非关节指令,让通用人形策略直接使用人类动作数据训练。其训练演示池含 1.406 亿帧,其中 93.2% 来自人类数据,在真机上零样本跟随移动指令成功率达 100%,而 GR00T N1.7 和 Ψ₀ 分别仅 16.7% 和 0%;无需任务专属微调即可达到 88.6% 的操作成功率,并可适配两种 VLA 与一种 world-action 模型骨干。
正文
Abstract:Teaching a humanoid to follow instructions with its whole body runs into two obstacles. Its action space is large and tightly coupled: legs, arms, and fingers must move together while the robot keeps its balance, which makes joint-level actions hard to learn. And humanoid demonstrations are scarce, so current humanoid generalist policies do not follow new instructions out of the box and are fine-tuned on teleoperated demonstrations of each task before deployment. Human demonstrations exist in far larger numbers, but a person's motion is not a robot command. We remove both obstacles by changing what the generalist policy predicts. We introduce VioLA, a generalist humanoid policy that predicts body and hand motion latents instead of joint commands. A pretrained body- and hand-controller execute these latents on the robot. Their corresponding motion encoders map human and robot motion into the same latent spaces. A human recording is therefore labeled in the policy's action space, and the training demonstration pool contains 140.6 million frames, 93.2% of them human. As a result, VioLA follows locomotion instructions on the real robot zero-shot, without task-specific fine-tuning, reaching 100% success where GR00T N1.7 and $\Psi_0$ reach 16.7% and 0%, respectively. It also reaches 88.6% manipulation success without task-specific fine-tuning. The same approach works across two VLA and one world-action model backbones. A generalist policy trained on human demonstrations alone performs locomotion tasks on the real robot zero-shot. Code and checkpoints will be released.
| Subjects: | Robotics (cs.RO); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.12435 [cs.RO] |
| (or arXiv:2610.12435v1 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12435 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mert Albaba [view email]
[v1]
Thu, 8 Oct 2026 17:57:27 UTC (6,440 KB)
来源:arXiv:cs.LG · arxiv.org