arXiv:cs.AI· Fei Zhang, Zhaochong An, Duncan Frost, Yikai Wang, Pengfei Liu, Ya Zhang, Michal Drozdzal, Amir Bar·· 5 小时前AI 评分45
ProWAM:用渐进式视觉规划实现世界动作建模
World Action Modeling with Progressive Visual Planning
AI 导读
ProWAM 是一种渐进式世界动作模型,通过联合预测动作与有序稀疏视觉子目标序列,为动作生成提供显式视觉引导,并可从大规模无动作视频中学习子目标预测。
正文
Abstract:World action models (WAMs) have emerged as a promising paradigm for robotic control by jointly predicting future visual dynamics and actions from an initial observation and instruction. However, existing WAMs struggle with long-horizon prediction, as generating dense video rollouts is highly inefficient. Some recent WAMs address this by predicting a single future frame without generating the full video, but this approach neglects how to progress toward the goal. We present ProWAM, a progressive world action model that jointly predicts actions and an ordered sequence of sparse visual sub-goals, providing explicit visual guidance to anchor action generation throughout task execution. This design scales naturally, as sub-goal prediction can be learned from large-scale action-free videos, allowing the video backbone to offload complex visual planning from the action policy. For efficient action generation, ProWAM executes a single video-backbone forward pass to cache sparse sub-goal features, eliminating iterative full-video generation and requiring only lightweight action denoising during replanning. Across extensive evaluations, ProWAM achieves superior out-of-distribution robustness. On simulation benchmarks, it sets new state-of-the-art results on LIBERO-Plus (85.8%) and randomized RoboTwin (75.7%), outperforming the strongest baseline with relative gains of up to +35.9%. On RoboCasa365, ProWAM achieves a 48.1% success rate and 18.2% on the challenging Composite-Unseen split, ranking 4th overall. Crucially, in zero-shot real-world experiments, ProWAM achieves 70.0% success, outperforming the strongest baseline by +15.0 (from 55.0% to 70.0%, a +27.3% relative gain) in novel scenes. These results demonstrate the value of progress-indexed visual foresight for closed-loop control. Our program is in this https URL.
| Comments: | Project Page: this https URL |
| Subjects: | Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO) |
| Cite as: | arXiv:2610.02508 [cs.AI] |
| (or arXiv:2610.02508v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02508 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Fei Zhang [view email]
[v1]
Thu, 1 Oct 2026 21:32:29 UTC (9,696 KB)
来源:arXiv:cs.AI · arxiv.org