arXiv:cs.LG· Liheng Ma, Rui Heng Yang, Amin Abyaneh, George Z. Xue, Behnam Rahmati, Mateo Clemente, Ziwen Hu, Anlin Chen, Tongtong Cao, Zhanguang Zhang, Yingxue Zhang·· 6 小时前AI 评分45
Faster-WAM:世界动作模型需要深层动作模块吗?
Faster-WAM: Do World Action Models Need Deep Action Modules?
AI 导读
研究提出以世界模型为中心的 Faster-WAM,仅用单层轻量动作专家,通过 DoT 与 Lite KV-Fusion、仅基于世界模型的条件化以及 retracted 1D-RoPE,把容量集中在视频世界模型中。
正文
Authors:Liheng Ma, Rui Heng Yang, Amin Abyaneh, George Z. Xue, Behnam Rahmati, Mateo Clemente, Ziwen Hu, Anlin Chen, Tongtong Cao, Zhanguang Zhang, Yingxue Zhang
Abstract:World Action Models (WAMs) build on pretrained video models, whose representations are grounded in physical dynamics and provide a natural basis for action prediction. Despite this natural foundation, many WAMs still rely on deep, parameter-heavy action-prediction modules that incur high inference latency and may overfit to limited robot demonstrations, restricting their real-world applicability. In this paper, we advocate a world-model-centric principle that concentrates capacity and computation in the video world model, while a lightweight action expert translates the backbone's representations into executable robot actions. We realize this principle through three key choices: Dock of Transformers (DoT) with Lite KV-Fusion to give the shallow, lightweight action expert access to representations from all video layers; world-model-only conditioning of the action expert; and retracted 1D-RoPE for positional alignment between video keys and action queries. We test this principle using Faster-WAM, a world-model-centric WAM with only a single-layer action expert. Despite this restriction on action-specific computation, Faster-WAM achieves competitive control performance on LIBERO and RoboTwin~2.0 without additional embodied pretraining. It provides approximately $3.7\times$ and $1.3\times$ inference speedups over Fast-WAM and $\pi_{0.5}$, respectively. Consistent with its world-model-centric design, Faster-WAM demonstrates stronger generalizability under distribution shifts: the same LIBERO-trained policy achieves $78.3\%$ success on LIBERO-Plus, exceeding Fast-WAM and LingBot-VA by $26.8$ and $8.8$ percentage points, respectively. Finally, real-robot experiments demonstrate success rates comparable to Fast-WAM, with substantially lower inference latency and shorter task-completion times.
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO) |
| Cite as: | arXiv:2608.02365 [cs.AI] |
| (or arXiv:2608.02365v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2608.02365 arXiv-issued DOI via DataCite |
Submission history
From: Liheng Ma [view email]
[v1]
Mon, 3 Aug 2026 15:11:21 UTC (493 KB)
[v2]
Tue, 6 Oct 2026 18:23:15 UTC (12,101 KB)
来源:arXiv:cs.LG · arxiv.org