跳到正文
arXiv:cs.AI· Qi Lyu, Jiahua Dong, Hao Shen, Xudong Wang, Hongyuan Yu, Baichen Liu, Henghui Ding, Zhi Han, Nicu Sebe, Ivan Laptev, Fahad Shahbaz Khan, Salman Khan·· 4 小时前AI 评分35

从历史动作轨迹中学习技能:面向世界动作模型的行动经验字典

Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models

AI 导读

研究者提出行动经验字典(AED),将历史物理动作轨迹编码为共享动作嵌入,以支持世界动作模型(WAMs)的技能复用与跨任务关系建模。该方法用预训练动作 tokenizer 从 AED 检索动作嵌入,经 cross-attention 视觉条件化后前置到噪声动作 token,并引入 motion-aware transition loss 监督随机时间间隔的视觉特征变化预测。

正文

Authors:Qi Lyu, Jiahua Dong, Hao Shen, Xudong Wang, Hongyuan Yu, Baichen Liu, Henghui Ding, Zhi Han, Nicu Sebe, Ivan Laptev, Fahad Shahbaz Khan, Salman Khan

View PDF HTML (experimental)

Abstract:World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redundant background elements interfere with the extraction of key visual information. To address these challenges, we develop a novel Action Experience Dictionary (AED) that encodes historical physical action trajectories into shared action embeddings to support skill reuse and model cross-task relationships. Specifically, we first aggregate historical actions to align with visual observations and retrieve action embeddings from the AED using a pretrained action tokenizer. Subsequently, we visually condition the pooled embeddings through cross-attention and prepend them to noisy action tokens, providing interaction context and action intent for prediction. To model action-related motion and reduce reliance on irrelevant background cues, we introduce a motion-aware transition loss that supervises visual feature change prediction over random temporal intervals. Experiments on simulation benchmarks and in real-world cross-embodiment settings verify the effectiveness of our AED. The project code is available at this https URL .
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.40219 [cs.CV]
  (or arXiv:2609.40219v2 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.40219

arXiv-issued DOI via DataCite

Submission history

From: Qi Lyu [view email]
[v1] Wed, 30 Sep 2026 17:26:38 UTC (21,279 KB)
[v2] Fri, 2 Oct 2026 05:19:23 UTC (21,278 KB)

来源:arXiv:cs.AI · arxiv.org