arXiv:cs.AI· Linhan Wang, Zijian An, Mingyuan Zhang, Chen Dai, Yi Xu, Can Cui, Jiayan Wang, Zichong Yang, Yinlin Chen, Lifeng Zhou, Chang-Tien Lu·· 5 小时前AI 评分47
GlanceWAM:面向世界-动作模型的稀疏测试时想象
GlanceWAM: Sparse Test-Time Imagination for World-Action Models
AI 导读
GlanceWAM 将视觉想象从控制回路中解耦,在共享视频 DiT 主干上以异步方式提前想象未来单帧,动作头则纯在潜空间中按控制频率解码动作块,单块延迟 48 ms(单张 A100),较同步世界-动作模型降低 24 倍。
正文
Authors:Linhan Wang, Zijian An, Mingyuan Zhang, Chen Dai, Yi Xu, Can Cui, Jiayan Wang, Zichong Yang, Yinlin Chen, Lifeng Zhou, Chang-Tien Lu
Abstract:Video generative models provide rich physical priors for robot learning, yet existing world-action models (WAMs) face a fundamental trade-off: synchronous video generation at control rate is latency-prohibitive, while abandoning test-time visual imagination sacrifices task success. We show that visual imagination achieves both real-time inference and superior success rates when generated asynchronously off the critical path and consumed directly in latent space. We introduce GlanceWAM, which decouples imagination from control on a single shared video DiT backbone: an asynchronous proposer glances ahead on a slow clock to imagine a single lookahead frame seconds into the future in the background, while an action head decodes action chunks at control rate (48 ms) purely in latent space without blocking. Enabled by a non-interfering attention mask that isolates video representations and staleness-robust horizon training that accommodates asynchronous lookahead aging, GlanceWAM breaks the speed-success dilemma. Trained purely on demonstrations, it attains 72.2% on the 24-task RoboCasa kitchen benchmark (vs. 67.1% for synchronous Cosmos Policy) and 99.0% on LIBERO while cutting per-chunk control latency $24\times$ relative to synchronous world-action models (48 ms on one A100). In single-arm and bimanual real-robot manipulation, it achieves higher average success than $\pi_{0.5}$ without any robot-data pretraining. Code is available at this https URL.
| Comments: | Add real-robot experiments |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO) |
| Cite as: | arXiv:2608.23927 [cs.CV] |
| (or arXiv:2608.23927v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2608.23927 arXiv-issued DOI via DataCite |
Submission history
From: Linhan Wang [view email]
[v1]
Tue, 25 Aug 2026 00:19:09 UTC (4,475 KB)
[v2]
Tue, 29 Sep 2026 19:01:04 UTC (10,056 KB)
来源:arXiv:cs.AI · arxiv.org