跳到正文
arXiv:cs.LG· Tingting Du, Ziyao Wang, Guoheng Sun, Ang Li·· 6 小时前AI 评分43

XGenAct:通过跨任务生成实现几何增强的世界动作模型

XGenAct: Geometry-Enhanced World Action Models through Cross-Task Generation

AI 导读

XGenAct 是一种世界动作模型,通过确定性编解码器将 RGB 观测、机器人动作、度量深度、表面法线和功能角色分割统一表示为 RGB 视频,并用单个视频扩散 Transformer 和一个目标函数跨这些空间学习时序预测,无需模态专用学习头。

正文

View PDF HTML (experimental)

Abstract:World action models (WAMs) have advanced robot control by predicting how observations and actions evolve over time. Despite this progress, RGB and action based future prediction does not explicitly address the spatial understanding needed for robot manipulation. Existing efforts often add a limited set of spatial prediction tasks through specialized heads or branches, leaving both the range of spatial supervision and the model architecture fragmented. We introduce XGenAct, a world action model that represents RGB observations, robot actions, metric depth, surface normals, and functional role segmentation as RGB videos through deterministic codecs. By sampling perception and action tasks during training, XGenAct uses one video diffusion transformer and one objective to learn temporal prediction across these spaces without modality specific learned heads. On held out RLBench tasks, structured perception training improves average closed loop success over RGB only training, and XGenAct achieves 52% success in the five task external comparison, versus 26% for the strongest evaluated baselines. It also predicts future depth and segmentation more accurately than the evaluated pipelines that generate RGB first and then apply a frozen perception expert.
Comments: 27 pages, including appendix
Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as: arXiv:2610.03516 [cs.RO]
  (or arXiv:2610.03516v1 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2610.03516

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Tingting Du [view email]
[v1] Fri, 2 Oct 2026 16:10:53 UTC (9,657 KB)

来源:arXiv:cs.LG · arxiv.org