arXiv:cs.AI· Dorian Benhamou Goldfajn, Mason Nakamura, Saaduddin Mahmud, Justin Svegliato, Kyle H. Wray, Shlomo Zilberstein·· 3 小时前
RoboTalk:从多模态演示中学习多机器人通信与协调
RoboTalk: Learning Multi-Robot Communication and Coordination from Multimodal Demonstrations
AI 导读
RoboTalk 是一个合成数据生成管线与数据集,包含 7,950 条多模态轨迹、覆盖 53 个移动操作厨房任务,用于训练小型 VLM 进行通信与协调。数据集包含 leader-follower 规划协议、工具调用(感知、操作、导航、通信)、推理轨迹与多样化自然语言通信。开源模型在其上微调后在新任务上达到 77% 成功率,而未微调的开源模型成功率仅约 2%。
正文
Abstract:Multi-robot collaboration could enable more efficient and scalable solutions to complex robotic tasks, but collaboration under partial observability remains challenging. Natural-language communication offers a promising approach to coordinating robots under partial observability. However, in decentralized manipulation, jointly learning explicit inter-robot communication and skill-level action selection from multimodal demonstrations remains underexplored for small vision-language models (VLMs) intended for on-device deployment. To address this gap, we introduce RoboTalk, a synthetic data-generation pipeline and dataset of 7,950 multimodal trajectories spanning 53 mobile-manipulation kitchen tasks for training small VLMs to communicate and coordinate. The dataset includes a leader-follower planning protocol, tool calls (perception, manipulation, navigation, and communication), rationale traces, and diversified natural-language communication. Fine-tuning open-source models on our dataset can reach 77% success on novel held-out tasks, a significant improvement over the untuned open source models, which had a success rate of around ~2%.
| Subjects: | Robotics (cs.RO); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.23997 [cs.RO] |
| (or arXiv:2609.23997v2 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2609.23997 arXiv-issued DOI via DataCite |
Submission history
From: Dorian Benhamou Goldfajn [view email]
[v1]
Mon, 21 Sep 2026 02:01:04 UTC (4,358 KB)
[v2]
Thu, 8 Oct 2026 16:19:43 UTC (4,358 KB)
来源:arXiv:cs.AI · arxiv.org