跳到正文
arXiv:cs.LG· Junwon Seo, Andrea Bajcsy·· 4 小时前AI 评分36

世界模型潜在扰动建模如何实现鲁棒决策

Modeling Latent Disturbances for Robust Decision-Making in World Models

AI 导读

研究提出将潜在空间扰动建模为对世界模型所学潜在动力学的扰动,通过结合动力学感知相似度度量与分布外检测构建合理潜在动力学集合,并用保形预测校准不确定性集,再以博弈论优化联合求解鲁棒动作与最坏情况扰动。在 Franka 机械臂实验中,该方法使安全过滤的失败率降低 70%,基于采样的策略引导失败率降低 54%。

正文

View PDF HTML (experimental)

Abstract:In this paper, we study robust decision-making in the latent space of world models (WMs). Robust optimization is a mathematical framework where, given explicitly specified dynamics and physically meaningful disturbances, a robot can select actions that remain effective even under worst-case disturbances. However, applying this principle to the learned latent space of WMs introduces a fundamental challenge: because WMs have fully learned state spaces and dynamics inferred from high-dimensional observations, it is unclear how to define latent-space disturbances that faithfully represent uncertainty in the underlying system. Our key idea is to model a latent-space disturbance as a perturbation to the learned latent dynamics that induces pessimistic but plausible transitions. Specifically, we construct a set of plausible latent dynamics by combining a dynamics-aware similarity metric that captures plausible transitions with out-of-distribution detection that excludes implausible latent states. We calibrate this uncertainty set over latent dynamics using conformal prediction, ensuring that WM imaginations induced by the latent disturbance remain plausible without becoming overly pessimistic. We then jointly optimize robust robot actions and the worst-case latent disturbances through game-theoretic optimization. We leverage this latent-space robust optimization to robustify policy steering, considering two paradigms: latent safety filtering and sample-and-verify steering of a generative control policy. Our controlled simulation experiments show that our latent disturbance enables robust decision-making directly in WM latent spaces, and hardware experiments with a Franka manipulator show that modeling latent disturbances enables robust policy steering, reducing failures by 70% in safety filtering and 54% in sampling-based policy steering. Project website: this https URL.
Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.07599 [cs.RO]
  (or arXiv:2610.07599v1 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2610.07599

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Junwon Seo [view email]
[v1] Tue, 6 Oct 2026 01:38:38 UTC (2,456 KB)

来源:arXiv:cs.LG · arxiv.org