跳到正文
arXiv:cs.LG· Liheng Ma, Rui Heng Yang, Amin Abyaneh, George Z. Xue, Behnam Rahmati, Mateo Clemente, Ziwen Hu, Anlin Chen, Tongtong Cao, Zhanguang Zhang, Yingxue Zhang·· 6 小时前AI 评分45

Faster-WAM:世界动作模型需要深层动作模块吗?

Faster-WAM: Do World Action Models Need Deep Action Modules?

AI 导读

研究提出以世界模型为中心的 Faster-WAM,仅用单层轻量动作专家,通过 DoT 与 Lite KV-Fusion、仅基于世界模型的条件化以及 retracted 1D-RoPE,把容量集中在视频世界模型中。

正文

Authors:Liheng Ma, Rui Heng Yang, Amin Abyaneh, George Z. Xue, Behnam Rahmati, Mateo Clemente, Ziwen Hu, Anlin Chen, Tongtong Cao, Zhanguang Zhang, Yingxue Zhang

View PDF HTML (experimental)

Abstract:World Action Models (WAMs) build on pretrained video models, whose representations are grounded in physical dynamics and provide a natural basis for action prediction. Despite this natural foundation, many WAMs still rely on deep, parameter-heavy action-prediction modules that incur high inference latency and may overfit to limited robot demonstrations, restricting their real-world applicability. In this paper, we advocate a world-model-centric principle that concentrates capacity and computation in the video world model, while a lightweight action expert translates the backbone's representations into executable robot actions. We realize this principle through three key choices: Dock of Transformers (DoT) with Lite KV-Fusion to give the shallow, lightweight action expert access to representations from all video layers; world-model-only conditioning of the action expert; and retracted 1D-RoPE for positional alignment between video keys and action queries. We test this principle using Faster-WAM, a world-model-centric WAM with only a single-layer action expert. Despite this restriction on action-specific computation, Faster-WAM achieves competitive control performance on LIBERO and RoboTwin~2.0 without additional embodied pretraining. It provides approximately $3.7\times$ and $1.3\times$ inference speedups over Fast-WAM and $\pi_{0.5}$, respectively. Consistent with its world-model-centric design, Faster-WAM demonstrates stronger generalizability under distribution shifts: the same LIBERO-trained policy achieves $78.3\%$ success on LIBERO-Plus, exceeding Fast-WAM and LingBot-VA by $26.8$ and $8.8$ percentage points, respectively. Finally, real-robot experiments demonstrate success rates comparable to Fast-WAM, with substantially lower inference latency and shorter task-completion times.
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO)
Cite as: arXiv:2608.02365 [cs.AI]
  (or arXiv:2608.02365v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2608.02365

arXiv-issued DOI via DataCite

Submission history

From: Liheng Ma [view email]
[v1] Mon, 3 Aug 2026 15:11:21 UTC (493 KB)
[v2] Tue, 6 Oct 2026 18:23:15 UTC (12,101 KB)

来源:arXiv:cs.LG · arxiv.org