arXiv:cs.LG· Xinyu Zhang, Zhengtong Xu, Yutian Tao, Yeping Wang, Yu She, Abdeslam Boularias·· 4 小时前AI 评分49
基于视觉特征的世界模型新方法 RLA-WM:从 DINO 残差学习 Residual Latent Action
Learning Visual Feature-Based World Models via Residual Latent Action
AI 导读
研究者提出 Residual Latent Action(RLA),可从 DINO 残差中学习,具备可预测性、泛化性并编码时间进程。基于 RLA 的 RLA World Model(RLA-WM)通过 flow matching 预测 RLA 值,在仿真和真实数据集上超越 SOTA 的特征基与视频扩散世界模型,且比视频扩散快数个数量级。
正文
Abstract:World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature-based world models, on the other hand, predict future visual features instead of raw video pixels, offering a promising alternative that is more efficient and less prone to hallucination. However, current feature-based approaches rely on direct regression, which leads to blurry or collapsed predictions in complex interactions, while generative modeling in high-dimensional feature spaces still remains challenging. In this work, we discover that a new type of latent action representation, which we refer to as Residual Latent Action (RLA), can be easily learned from DINO residuals. We also show that RLA is predictive, generalizable, and encodes temporal progression. Building on RLA, we propose RLA World Model (RLA-WM), which predicts RLA values via flow matching. RLA-WM outperforms both state-of-the-art feature-based and video-diffusion world models on simulation and real-world datasets, while being orders of magnitude faster than video diffusion. Furthermore, we develop two robot learning techniques that use RLA-WM to improve policy learning. The first one is a minimalist world action model with RLA that learns from actionless videos, and improves VLA on LIBERO and real robot. The second one is a visual RL framework trained entirely inside a world model learned from offline videos only, using a video-aligned reward and no online interactions. Project page: this https URL
| Comments: | NeurIPS 2026 |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO) |
| Cite as: | arXiv:2605.07079 [cs.CV] |
| (or arXiv:2605.07079v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2605.07079 arXiv-issued DOI via DataCite |
Submission history
From: Xinyu Zhang [view email]
[v1]
Fri, 8 May 2026 00:58:16 UTC (19,939 KB)
[v2]
Mon, 5 Oct 2026 20:17:03 UTC (31,077 KB)
来源:arXiv:cs.LG · arxiv.org