arXiv:cs.LG· Rohun Agrawal, Nimit Kalra, Arjun Parthasarathy, Yann LeCun, Oumayma Bounou, Pavel Izmailov, Micah Goldblum·· 5 小时前AI 评分36
世界模型如何弥合训练-测试差距以实现基于梯度的规划
Closing the Train-Test Gap in World Models for Gradient-Based Planning
AI 导读
针对世界模型训练时用下一状态预测目标、测试时却要估计动作序列的错位,研究者提出训练时数据合成方法,让现有世界模型支持高效的基于梯度规划。在多种物体操作与导航任务上,该方法在仅用 10% 时间预算的情况下,表现超过或持平经典无梯度的交叉熵方法(CEM)。
正文
Abstract:World models paired with model predictive control (MPC) can be trained offline on large-scale datasets of expert trajectories and enable generalization to a wide range of planning tasks at inference time. Compared to traditional MPC procedures, which rely on slow search algorithms or on iteratively solving optimization problems exactly, gradient-based planning offers a computationally efficient alternative. However, the performance of gradient-based planning has thus far lagged behind that of other approaches. In this paper, we propose improved methods for training world models that enable efficient gradient-based planning. We begin with the observation that although a world model is trained on a next-state prediction objective, it is used at test-time to instead estimate a sequence of actions. The goal of our work is to close this train-test gap. To that end, we propose train-time data synthesis techniques that enable significantly improved gradient-based planning with existing world models. At test time, our approach outperforms or matches the classical gradient-free cross-entropy method (CEM) across a variety of object manipulation and navigation tasks in 10% of the time budget.
| Subjects: | Machine Learning (cs.LG); Robotics (cs.RO) |
| Cite as: | arXiv:2512.09929 [cs.LG] |
| (or arXiv:2512.09929v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2512.09929 arXiv-issued DOI via DataCite |
Submission history
From: Rohun Agrawal [view email]
[v1]
Wed, 10 Dec 2025 18:59:45 UTC (7,603 KB)
[v2]
Thu, 1 Oct 2026 18:46:00 UTC (12,753 KB)
来源:arXiv:cs.LG · arxiv.org