arXiv:cs.AI· Tianbin Liu, Jian Zhu, Taiyi Su, Jianjun Zhang, Chong Ma, Zitai Huang, Weiyi Lu, Yi Xu·· 6 小时前AI 评分44
CSWAM:基于 V-JEPA 2.1 因果语义专家提升世界动作模型的分布外泛化能力
CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models
AI 导读
研究者提出 Causal Semantic World Action Model(CSWAM),在 FastWAM 上引入基于 V-JEPA 2.1 的因果语义专家,学习语义状态变化与运动的未来演化,并通过因果注意力与视频、动作流共享历史上下文,推理时仍保持仅动作的高效推理。
正文
Abstract:FastWAM-style world action models enable efficient action-only inference, but generalize poorly under visual distribution shifts. Their reconstruction-oriented representations emphasize appearance-specific details, limiting generalization to unseen scenes and objects. Without observation history, the model also lacks temporal evidence for robustly identifying task-relevant state changes and motion in unfamiliar visual conditions. To address these limitations, we present the Causal Semantic World Action Model (CSWAM), which augments FastWAM with a causal semantic expert built on V-JEPA 2.1. V-JEPA provides temporally grounded representations of semantic state changes and motion with less dependence on appearance-specific details. The expert learns their future evolution from a sparse history of current and past observations and shares the history-derived context with both the video and action streams through causal attention. At inference, CSWAM conditions action denoising on the current video state and observed semantic history, retaining efficient action-only inference. We conduct simulation and real-robot experiments to evaluate generalization under distribution shifts. With embodied pretraining, CSWAM raises Randomized success on RoboTwin 2.0 Clean-to-Randomized transfer from 10.16% to 45.18%, a gain of 35.02 percentage points over FastWAM. Across two real-robot tasks and three OOD difficulty levels, CSWAM improves average success over FastWAM by 42.5 percentage points, from 27.5% to 70.0%.
| Comments: | 13 pages, 2 figures |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.18462 [cs.CV] |
| (or arXiv:2609.18462v5 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.18462 arXiv-issued DOI via DataCite |
Submission history
From: Jian Zhu [view email]
[v1]
Wed, 16 Sep 2026 10:56:39 UTC (1,172 KB)
[v2]
Thu, 17 Sep 2026 01:35:06 UTC (1,172 KB)
[v3]
Sun, 20 Sep 2026 02:41:43 UTC (961 KB)
[v4]
Sun, 4 Oct 2026 10:28:10 UTC (1,647 KB)
[v5]
Tue, 6 Oct 2026 10:28:06 UTC (1,543 KB)
来源:arXiv:cs.AI · arxiv.org