arXiv:cs.LG(机器学习,全量分类)· Jiaming Wang, Duc Thang Nguyen, Jizhuo Chen, Volodymyr Shcherbyna, Diwen Liu, Zhengcheng Shen, Harold Soh·· 5 小时前AI 评分45
从互联网视频学习社会导航:在策略状态空间中构建训练环境
Learning Social Navigation from Internet Videos in the Policy State Space
AI 导读
研究者提出一条高效流水线,将普通单目行走视频直接转换为策略状态空间中的闭环社会导航训练环境,用可通行性地图表示静态场景、直接回放视频中恢复的行人轨迹。所得策略在独立 Arena 基准上成功率达 81.2%,最强基线为 75.0%,并在 20 次真机试验中成功 19 次,无需策略微调。
正文
Abstract:Training robust social-navigation policies requires simulators with diverse scene layouts, terrain, and human motion, but constructing such environments and specifying pedestrian behavior is costly. We propose an efficient pipeline that converts ordinary monocular walking videos directly into closed-loop social-navigation training environments in the policy's state space. Our key observation is that local social navigation primarily depends on two types of information: where the robot can traverse and how nearby pedestrians move. We therefore represent the static scene as a metric traversability map, which can be rigidly transformed under counterfactual robot motion, while directly replaying the pedestrian trajectories recovered from the video over time. This abstraction allows us to define the forward dynamics directly in the policy's state space and efficiently simulate counterfactual robot states without reconstructing or rendering photorealistic observations. The resulting policy achieves 81.2% success in the independent Arena benchmark, compared with 75.0% for the strongest baseline, and succeeds in 19/20 real-robot trials without policy fine-tuning. Project page: this https URL
| Comments: | 9 pages, 5 figures, 6 tables |
| Subjects: | Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.37476 [cs.RO] |
| (or arXiv:2609.37476v2 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2609.37476 arXiv-issued DOI via DataCite |
Submission history
From: Jiaming Wang [view email]
[v1]
Sat, 26 Sep 2026 07:16:56 UTC (5,057 KB)
[v2]
Thu, 1 Oct 2026 09:12:24 UTC (5,058 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org