arXiv:cs.LG· Junwon Seo, Sushant Veer, Ran Tian, Wenhao Ding, Apoorva Sharma, Karen Leung, Edward Schmerling, Marco Pavone, Andrea Bajcsy·· 5 小时前AI 评分40
StressDream:引导视频世界模型实现稳健的策略评估与改进
StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement
AI 导读
StressDream 通过优化扩散式视频世界模型的初始噪声,将想象引导至推理时指定的高影响但合理的未来结果,如任务失败。该方法结合视觉语言模型提供语义梯度与合理性目标防止噪声分布外漂移,在自动驾驶和机器人操作的 SOTA 视频世界模型上验证有效,可识别那些合理未来包含不良后果的动作,从而实现稳健的策略评估与改进。
正文
Abstract:Video world models (WMs) have shown promise for policy evaluation and improvement by imagining realistic future observations conditioned on ego-robot actions. While WMs can model distributions over futures, policy evaluation and improvement typically rely on nominal imaginations, which can miss high-impact outcomes of robot actions unless prohibitively many samples are drawn. To enable robust policy evaluation and improvement over WM imaginations, we propose StressDream, which steers imaginations toward high-impact yet plausible outcomes specified at inference time by optimizing the initial noise of diffusion-based WMs. However, optimizing high-dimensional noise is challenging: the optimization must reason about nuanced, scene-dependent target events in generated videos while avoiding out-of-distribution (OOD) noise that yields implausible imaginations. We address this with two complementary objectives: a semantic objective with a Vision-Language Model that provides informative gradients by reasoning about the generated video, and a plausibility objective that prevents the optimized noise from drifting OOD. With state-of-the-art video world models for autonomous driving and robotic manipulation, we show that StressDream effectively steers imaginations toward high-impact yet plausible outcomes specified by text at inference time, such as task failures, enabling robust policy evaluation and improvement by identifying actions whose plausible futures include undesirable outcomes. Video results are available at this https URL.
| Comments: | Conference on Robot Learning (CoRL) 2026. Project page: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO) |
| Cite as: | arXiv:2606.00267 [cs.CV] |
| (or arXiv:2606.00267v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2606.00267 arXiv-issued DOI via DataCite |
Submission history
From: Junwon Seo [view email]
[v1]
Fri, 29 May 2026 18:57:57 UTC (8,514 KB)
[v2]
Tue, 6 Oct 2026 21:57:34 UTC (15,737 KB)
来源:arXiv:cs.LG · arxiv.org