跳到正文
arXiv:cs.LG· Samuel Barbeau, Simon Roy, Giovanni Beltrame, Christian Desrosiers, Nicolas Thome·· 5 小时前AI 评分38

LAGO:从语言预测潜在目标的分层世界模型规划方法

Latent Goal Prediction from Language for Model-Based Planning

AI 导读

研究者提出 LAGO(Latent Goal Prediction from Language),一种分层世界模型,用单一预测器同时预测动作条件动态并将语言指令落地为中间潜在子目标序列,两种模式共享同一潜在空间和回归目标。

正文

View PDF HTML (experimental)

Abstract:Joint-Embedding Predictive Architectures (JEPAs) enable agents to plan in latent space by imagining the outcomes of candidate actions, yet task specification remains a bottleneck. Visual targets provide precise local gradients but poor distant guidance, while language is flexible yet limited by noisy cross-modal alignment or dependence on distinct large generative models. We introduce LAGO (Latent Goal Prediction from Language), a hierarchical world model in which a single predictor both forecasts action-conditioned dynamics and grounds language instructions as sequences of intermediate latent subgoals, training both modes with a single regression objective over a shared latent space. At each planning step, LAGO predicts a sequence of latent subgoals from a language instruction and optimizes an action sequence using a soft-minimum alignment cost that rewards subgoal proximity without enforcing a rigid path. Subgoals are repredicted as the agent's state evolves, turning a single long-horizon instruction into a sequence of locally tractable objectives. Across environments spanning navigation and manipulation, LAGO outperforms flat and hierarchical image-goal planners, with the largest gains on long-horizon navigation with curved, hazard-constrained paths, where it more than doubles the success rate of a goal-conditioned behavior cloning policy trained on the same demonstrations. Grounding language in the world model's own latent space also outperforms vision-language reward models, including one that plans with ground-truth dynamics.
Comments: 10 pages. Preprint under review
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2606.20627 [cs.AI]
  (or arXiv:2606.20627v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2606.20627

arXiv-issued DOI via DataCite

Submission history

From: Samuel Barbeau [view email]
[v1] Tue, 26 May 2026 20:04:33 UTC (13,573 KB)
[v2] Thu, 1 Oct 2026 18:22:31 UTC (14,315 KB)

来源:arXiv:cs.LG · arxiv.org