跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Nathan Tsoi, Michael J. Munje, Tejas Oberoi, Rishab Maheshwari, Pengen Zheng, Tanush Chauhan, Peter Stone, Joydeep Biswas·· 5 小时前AI 评分37

STARS:从时空动态到人机交互中的社会表征

STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction

AI 导读

研究者提出 SocialNav-SUB,一个用于评估 VLM 在真实社交机器人导航场景中场景理解能力的 VQA 数据集与基准,覆盖空间、时空与社会推理任务。实验显示表现最好的 VLM 虽能与人类答案有较高一致概率,但仍不及更简单的规则方法和人类共识基线,表明当前 VLM 在社交场景理解上存在关键缺口。该基准为社交机器人导航基础模型研究提供统一评估框架,论文入选 CoRL 2026。

正文

View PDF HTML (experimental)

Abstract:Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding, capabilities that align with the nuanced requirements of social robot navigation. However, it remains unclear whether VLMs can accurately understand complex social navigation scenes (e.g., inferring the spatial-temporal relations among agents and human intentions), which is essential for safe and socially compliant robot navigation. While some recent works have explored the use of VLMs in social robot navigation, no existing work systematically evaluates their ability to meet these necessary conditions. In this paper, we introduce the Social Navigation Scene Understanding Benchmark (SocialNav-SUB), a Visual Question Answering (VQA) dataset and benchmark designed to evaluate VLMs for scene understanding in real-world social robot navigation scenarios. SocialNav-SUB provides a unified framework for evaluating VLMs against human and rule-based baselines across VQA tasks requiring spatial, spatiotemporal, and social reasoning in social robot navigation. Through experiments with state-of-the-art VLMs, we find that while the best-performing VLM achieves an encouraging probability of agreeing with human answers, it still underperforms simpler rule-based approach and human consensus baselines, indicating critical gaps in social scene understanding of current VLMs. Our benchmark sets the stage for further research on foundation models for social robot navigation, offering a framework to explore how VLMs can be tailored to meet real-world social robot navigation needs. An overview of this paper along with the code and data can be found at this https URL.
Comments: Conference on Robot Learning (CoRL) 2026. First two authors contributed equally. Project site: this https URL
Subjects: Robotics (cs.RO); Machine Learning (cs.LG)
Cite as: arXiv:2609.40245 [cs.RO]
  (or arXiv:2609.40245v2 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2609.40245

arXiv-issued DOI via DataCite

Submission history

From: Michael Munje [view email]
[v1] Wed, 30 Sep 2026 17:35:08 UTC (3,458 KB)
[v2] Thu, 1 Oct 2026 02:36:39 UTC (3,458 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org