arXiv:cs.AI· Jai Bardhan, Josef Sivic, Vladimir Petrik·· 6 小时前AI 评分44
DepthWorld:面向机器人操作的 3D 世界模型
DepthWorld: 3D World Model for Robot Manipulation
AI 导读
DepthWorld 是一个基于 Stable Video Diffusion 的 3D 世界模型,通过空间潜变量分块同时预测多视角 RGB 与深度,且不改动预训练 VAE。
正文
Abstract:World models offer a data-driven alternative to traditional simulators for robotics, with applications spanning policy evaluation, improvement, and planning. All of these uses depend on faithful 3D geometry, yet current video-based world models are trained on RGB alone and produce rollouts that look correct frame-by-frame but do not compose into a consistent 3D world. Closing this gap requires progress on two fronts: large-scale 3D supervision for manipulation, and an architecture that can absorb it without disturbing strong pretrained video priors. We introduce a calibration pipeline that combines learned stereo depth with a joint factor graph, pooling all episodes collected from the same physical robot to recover its shared kinematic parameters alongside per-scene extrinsics. Applied to the DROID dataset, this yields DROID-3D, a calibrated 3D dataset providing dense metric depth and recalibrated multi-view extrinsics (achieving <0.7 px reprojection error on 90% of episodes for external cameras). We then train DepthWorld, a Stable Video Diffusion-based world model that jointly predicts multi-view RGB and depth via spatial latent tiling, leaving the pretrained Variational Autoencoder (VAE) unchanged. Depth supervision improves RGB prediction itself by +1.48 dB PSNR over an identical RGB-only baseline at equal training budget, while simultaneously yielding accurate metric depth for downstream geometric reasoning.
| Comments: | Accepted at the Conference on Robot Learning (CoRL) 2026. Project page: this https URL. 32 pages including supplementary material, 15 figures, 7 tables |
| Subjects: | Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2610.08780 [cs.RO] |
| (or arXiv:2610.08780v1 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08780 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jai Bardhan [view email]
[v1]
Tue, 6 Oct 2026 17:59:00 UTC (43,670 KB)
来源:arXiv:cs.AI · arxiv.org