arXiv:cs.AI· Jicong Ao, Shuhan Jiang, Yuling Zhong, Yanwen Liu, Yuhan Gao, Jiangyuan Zhao, Yang Zhang, Shiqiang Zhu, Chenjia Bai, Xuelong Li·· 10 小时前AI 评分45
SMART:通过大规模合成预训练实现零样本仿真到现实的关节物体操作
SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining
AI 导读
研究团队提出 SMART 系统,基于自研仿真平台 SMART-Sim 合成超 100 万条演示数据 SMART-Data,覆盖 44 种原子任务类型、5 种机器人设置和 2,507 个关节物体。在该数据上预训练的视觉-语言-动作(VLA)模型在仿真基准上表现具竞争力,并实现零样本 sim-to-real 迁移和真实世界关节物体操作的可扩展性能。
正文
Abstract:The ability to interact with articulated objects is essential for embodied intelligent systems, but collecting large-scale real-world demonstrations for these interactions remains challenging due to the precise contact and constraint-following motions involved. Although simulation provides a promising alternative, existing synthetic data efforts cover limited articulated-object categories, while general-purpose synthesis pipelines lack explicit designs for part-level semantics and articulation constraints, hindering agentic task generation and scalable synthesis of high-quality articulated-manipulation demonstrations. To bridge this gap, we introduce SMART, a scalable system leveraging large-scale Synthesized Manipulation demonstrations for ARTiculated-object manipulation. At its core, we develop SMART-Sim, a simulation platform with articulation-aware design that enables effective task generation and efficient demonstration collection. Building on SMART-Sim, we apply agentic task generation and design a scalable distributed synthesis system, using them to synthesize SMART-Data, comprising over 1M demonstrations across 44 atomic task types, 5 robot setups, and 2,507 articulated objects. The vision-language-action (VLA) model pretrained on SMART-Data shows competitive performance on simulation benchmarks and achieves zero-shot sim-to-real transfer and scalable performance in real-world articulated-object manipulation tasks. This highlights the potential of synthetic demonstrations in providing effective and scalable supervision for improving VLA model performance in contact-rich articulated-object manipulation.
| Comments: | Technical Report, 31 pages |
| Subjects: | Robotics (cs.RO); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07652 [cs.RO] |
| (or arXiv:2610.07652v1 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07652 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jicong Ao [view email]
[v1]
Tue, 6 Oct 2026 02:47:22 UTC (44,530 KB)
来源:arXiv:cs.AI · arxiv.org