跳到正文
arXiv:cs.AI· Linhan Wang, Zijian An, Mingyuan Zhang, Chen Dai, Yi Xu, Can Cui, Jiayan Wang, Zichong Yang, Yinlin Chen, Lifeng Zhou, Chang-Tien Lu·· 5 小时前AI 评分47

GlanceWAM:面向世界-动作模型的稀疏测试时想象

GlanceWAM: Sparse Test-Time Imagination for World-Action Models

AI 导读

GlanceWAM 将视觉想象从控制回路中解耦,在共享视频 DiT 主干上以异步方式提前想象未来单帧,动作头则纯在潜空间中按控制频率解码动作块,单块延迟 48 ms(单张 A100),较同步世界-动作模型降低 24 倍。

正文

Authors:Linhan Wang, Zijian An, Mingyuan Zhang, Chen Dai, Yi Xu, Can Cui, Jiayan Wang, Zichong Yang, Yinlin Chen, Lifeng Zhou, Chang-Tien Lu

View PDF HTML (experimental)

Abstract:Video generative models provide rich physical priors for robot learning, yet existing world-action models (WAMs) face a fundamental trade-off: synchronous video generation at control rate is latency-prohibitive, while abandoning test-time visual imagination sacrifices task success. We show that visual imagination achieves both real-time inference and superior success rates when generated asynchronously off the critical path and consumed directly in latent space. We introduce GlanceWAM, which decouples imagination from control on a single shared video DiT backbone: an asynchronous proposer glances ahead on a slow clock to imagine a single lookahead frame seconds into the future in the background, while an action head decodes action chunks at control rate (48 ms) purely in latent space without blocking. Enabled by a non-interfering attention mask that isolates video representations and staleness-robust horizon training that accommodates asynchronous lookahead aging, GlanceWAM breaks the speed-success dilemma. Trained purely on demonstrations, it attains 72.2% on the 24-task RoboCasa kitchen benchmark (vs. 67.1% for synchronous Cosmos Policy) and 99.0% on LIBERO while cutting per-chunk control latency $24\times$ relative to synchronous world-action models (48 ms on one A100). In single-arm and bimanual real-robot manipulation, it achieves higher average success than $\pi_{0.5}$ without any robot-data pretraining. Code is available at this https URL.
Comments: Add real-robot experiments
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)
Cite as: arXiv:2608.23927 [cs.CV]
  (or arXiv:2608.23927v2 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2608.23927

arXiv-issued DOI via DataCite

Submission history

From: Linhan Wang [view email]
[v1] Tue, 25 Aug 2026 00:19:09 UTC (4,475 KB)
[v2] Tue, 29 Sep 2026 19:01:04 UTC (10,056 KB)

来源:arXiv:cs.AI · arxiv.org