跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Takumi Hara, Kanata Suzuki·· 1 天前AI 评分31

监督决定成功的量:面向潜在世界模型规划的准则对齐辅助损失

Supervise What Decides Success: Criterion-Aligned Auxiliary Losses for Latent World-Model Planning

AI 导读

针对潜在世界模型以潜空间距离评分动作序列、却无法区分成功候选的问题,研究者提出一种准则对齐辅助损失:训练时用线性头从编码器与预测器输出回归成功准则物理量,并将回归误差加入损失,训练后丢弃该头,测试时模型、开销与输入均不变。该损失使 PushT 和 cube 任务成功率分别绝对提升 3.5% 和 3.4%,两项提升均具统计显著性。

正文

View PDF HTML (experimental)

Abstract:Latent world models plan by scoring candidate action sequences with distances in latent space. However, task success is judged by physical quantities, which we call the success-criterion quantities. In all four latent world models we examine, the end-effector position is encoded in the latent state with an error larger than the success criterion allows. Such a latent state cannot separate successful candidates from failing ones. We propose an auxiliary loss that uses success-criterion quantities as training targets, whereas existing latent world models take them only as inputs. During training, a linear head on the encoder and predictor outputs regresses the success-criterion quantities, and the regression error is added to the training loss. The head is discarded after training, so the model, its cost, and its inputs at test time are unchanged. This loss alone improves the success rate on PushT and cube by 3.5% and 3.4% (absolute), respectively, and both improvements are statistically significant. A success criterion thus specifies what a world model must retain in its latent state, and we show that it can serve directly as a training target.
Comments: 21 pages, 6 figures, 10 tables. Under review
Subjects: Machine Learning (cs.LG); Robotics (cs.RO)
Cite as: arXiv:2610.01224 [cs.LG]
  (or arXiv:2610.01224v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01224

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Takumi Hara [view email]
[v1] Thu, 1 Oct 2026 07:27:10 UTC (1,675 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org