跳到正文
arXiv:cs.AI· Yan Yang, Jikun Rong, Minzhao Zhu, Zheyi Zhao, Qirui Hu, Zihan Lan, Weixin Mao, Yinhao Li, Zhen Fu, Hua Chen·· 3 小时前

HWAM:联合状态-动作生成的人形世界动作模型

Humanoid World Action Model With Joint State--Action Generation

AI 导读

针对分层人形系统中参考动作与实际执行动作之间存在"动作-执行鸿沟"的问题,研究者提出 HWAM——一种联合生成参考动作与执行后本体状态的人形世界动作模型。该模型通过 Policy、前向动力学建模(FDM)和逆动力学建模(IDM)三条互补条件路径训练,将执行动作的监督直接纳入动作学习。

正文

View PDF HTML (experimental)

Abstract:Humanoid robots are a promising platform for general-purpose manipulation. Recent Vision-Language-Action (VLA) policies learn actions directly from multimodal observations, while World Action Models (WAMs) further incorporate future visual prediction to improve action generation. However, in hierarchical humanoid systems, VLA and WAM policies output reference actions that are subsequently realized through whole-body control, robot dynamics, balance, and contact. This hierarchy creates an action--execution gap: the reference produced by the policy can differ from the motion realized by the robot. Without explicitly modeling the realized body state, future visual prediction must jointly explain scene evolution and discrepancies between reference actions and executed motion, making it difficult to associate an action with its physical outcome. We propose HWAM, a Humanoid World Action Model with joint state--action generation, which makes the robot's post-execution proprioceptive state an explicit prediction target. By jointly generating reference actions and their realized body states, HWAM directly incorporates supervision of executed motion into action learning. HWAM is trained through three complementary conditional paths. The Policy path jointly denoises state--action trajectories conditioned only on current observations, matching deployment conditions. Forward Dynamics Modeling (FDM) predicts future visual observations conditioned on actions and post-execution states, while Inverse Dynamics Modeling (IDM) reconstructs the joint trajectory from visual transitions. Together, these paths connect policy references, realized body motion, and visual outcomes. HWAM achieves the highest success rate among evaluated baselines on three real-robot tasks on the LimX OLI humanoid. On Candy Picking, HWAM achieves a 70.6% success rate, compared with 43.3% for Fast-WAM.
Comments: Under review. 15 pages, 5 figures
Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.12026 [cs.RO]
  (or arXiv:2610.12026v1 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2610.12026

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yan Yang [view email]
[v1] Thu, 8 Oct 2026 14:23:18 UTC (10,574 KB)

来源:arXiv:cs.AI · arxiv.org