arXiv:cs.LG(机器学习,全量分类)· Deqian Kong, Guangyan Sun, Sheng Cheng, Sirui Xie, Bo Pang, Jianwen Xie, Tony Geng, Caiwen Ding, Ying Nian Wu·· 5 小时前AI 评分43
从随机探索中学习规划:基于条件能量模型的多尺度认知地图
Learning to Plan from Random Exploration
AI 导读
研究提出一种条件能量模型,通过噪声对比估计在无动作和奖励标签的观测对上训练,从随机探索中学习时间对数密度比,实现长程迷宫规划。该模型在不同时间跨度上查询学到的关系,配合局部动力学模型预测候选动作结果并评估目标进展,在状态和图像输入下均完成规划。实验显示学到的分数场、嵌入向量探测和规划路径呈现多尺度认知地图特性,并支持自我中心导航和次优数据下的操作规划。
正文
Abstract:Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without policy-improvement training? Our random-walk analysis explains what temporal relations contain: short horizons reveal geodesic geometry in the diffusion limit, while longer horizons reveal connectivity between regions before mixing removes these distinctions. We learn these relations with a conditional energy-based model that estimates temporal log-density ratios through horizon-conditioned embeddings. The model is trained on observation pairs by noise-contrastive estimation, without action or reward labels. The planner queries these learned relations at different horizons as it moves toward the goal. At test time, a separate local dynamics model predicts candidate action outcomes, and the temporal model evaluates their progress toward the goal by selecting or aggregating estimated improvements across horizons. The agent executes one action and replans with both models fixed. Experiments demonstrate long-range maze planning from random exploration using states and images. Learned score fields, embedding probes, and planned routes exhibit properties of a multiscale cognitive map. We further demonstrate egocentric navigation from random exploration and manipulation planning from suboptimal data.
| Subjects: | Machine Learning (cs.LG); Robotics (cs.RO); Machine Learning (stat.ML) |
| Cite as: | arXiv:2609.38383 [cs.LG] |
| (or arXiv:2609.38383v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38383 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Kong Deqian [view email]
[v1]
Tue, 29 Sep 2026 18:38:49 UTC (722 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org