跳到正文
arXiv:cs.AI· Tianbin Liu, Jian Zhu, Taiyi Su, Jianjun Zhang, Chong Ma, Zitai Huang, Weiyi Lu, Yi Xu·· 6 小时前AI 评分44

CSWAM:基于 V-JEPA 2.1 因果语义专家提升世界动作模型的分布外泛化能力

CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models

AI 导读

研究者提出 Causal Semantic World Action Model(CSWAM),在 FastWAM 上引入基于 V-JEPA 2.1 的因果语义专家,学习语义状态变化与运动的未来演化,并通过因果注意力与视频、动作流共享历史上下文,推理时仍保持仅动作的高效推理。

正文

View PDF HTML (experimental)

Abstract:FastWAM-style world action models enable efficient action-only inference, but generalize poorly under visual distribution shifts. Their reconstruction-oriented representations emphasize appearance-specific details, limiting generalization to unseen scenes and objects. Without observation history, the model also lacks temporal evidence for robustly identifying task-relevant state changes and motion in unfamiliar visual conditions. To address these limitations, we present the Causal Semantic World Action Model (CSWAM), which augments FastWAM with a causal semantic expert built on V-JEPA 2.1. V-JEPA provides temporally grounded representations of semantic state changes and motion with less dependence on appearance-specific details. The expert learns their future evolution from a sparse history of current and past observations and shares the history-derived context with both the video and action streams through causal attention. At inference, CSWAM conditions action denoising on the current video state and observed semantic history, retaining efficient action-only inference. We conduct simulation and real-robot experiments to evaluate generalization under distribution shifts. With embodied pretraining, CSWAM raises Randomized success on RoboTwin 2.0 Clean-to-Randomized transfer from 10.16% to 45.18%, a gain of 35.02 percentage points over FastWAM. Across two real-robot tasks and three OOD difficulty levels, CSWAM improves average success over FastWAM by 42.5 percentage points, from 27.5% to 70.0%.
Comments: 13 pages, 2 figures
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.18462 [cs.CV]
  (or arXiv:2609.18462v5 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.18462

arXiv-issued DOI via DataCite

Submission history

From: Jian Zhu [view email]
[v1] Wed, 16 Sep 2026 10:56:39 UTC (1,172 KB)
[v2] Thu, 17 Sep 2026 01:35:06 UTC (1,172 KB)
[v3] Sun, 20 Sep 2026 02:41:43 UTC (961 KB)
[v4] Sun, 4 Oct 2026 10:28:10 UTC (1,647 KB)
[v5] Tue, 6 Oct 2026 10:28:06 UTC (1,543 KB)

来源:arXiv:cs.AI · arxiv.org