arXiv:cs.AI· Florian Strohm, Patrick Wagner, Jannik Schwab, Marco Huber·· 5 小时前AI 评分46
JEPA 世界模型如何在画面几乎不动时保持可规划性
Keeping JEPA World Models Plannable When Little of the Frame Moves
AI 导读
研究者构建 SLIM 推物基准,用配对视觉与语言目标测试潜在世界模型的语言规划能力,发现 LeWM 世界模型成功率不足 1%,探针定位故障在编码器:其潜在表示几乎对动作不敏感。加入一项逆动力学辅助损失后,成功率从 0.003 升至 0.35,困难推物层级达 0.16,并在两倍训练视野下改善 PushT。修复后的潜在表示上,小型语言目标头无需重训世界模型即可从句子规划,导航达 0.84。
正文
Abstract:Specifying a goal in language rather than as a goal frame is a natural interface for planning with a latent world model, but testing it needs scenes in which language must discriminate between several objects. We build SLIM, a pushing benchmark with several small objects and paired visual and language goals on identical scenes. On SLIM a LeWM world model that solves PushT succeeds on under 1% of trials, although a scripted controller with simulator state solves every tier. Probes locate the failure in the encoder: its latent is nearly action-insensitive, neither pusher nor object positions can be decoded from it, and rollouts are no better than copying the current latent forward. One inverse-dynamics auxiliary loss, applied to encoder latents and to predicted latents through a shared head discarded at test time, restores every probe and raises success from 0.003 to 0.35 (0.16 on the hard pushing tier, where a goal-agnostic policy scores zero), and improves PushT at twice the trained horizon. Controls attribute the repair to the gradient into the encoder, and a response sweep shows that the vanilla model plans once enough of the frame responds to actions. A cheap action-sensitivity probe, computable without environment access, acts as an empirical necessary condition: all configurations below its threshold failed to plan. On the repaired latent, a small language-goal head plans from sentences without retraining the world model: it reaches 0.84 on navigation (visual-goal oracle 1.00), follows the named zone when it is swapped with a decoy, and degrades gracefully to unseen nouns. A single goal sentence rarely completes a push, but given the push as a sequence of stage sentences the head raises success on the medium and hard pushing tiers from 0.04 to 0.25, on par with the goal-frame oracle, also when the switch between stages is read from the latent alone.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.03137 [cs.AI] |
| (or arXiv:2610.03137v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03137 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Florian Strohm [view email]
[v1]
Fri, 2 Oct 2026 11:02:39 UTC (198 KB)
来源:arXiv:cs.AI · arxiv.org