跳到正文
arXiv:cs.AI· Jicong Ao, Shuhan Jiang, Yuling Zhong, Yanwen Liu, Yuhan Gao, Jiangyuan Zhao, Yang Zhang, Shiqiang Zhu, Chenjia Bai, Xuelong Li·· 10 小时前AI 评分45

SMART:通过大规模合成预训练实现零样本仿真到现实的关节物体操作

SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining

AI 导读

研究团队提出 SMART 系统,基于自研仿真平台 SMART-Sim 合成超 100 万条演示数据 SMART-Data,覆盖 44 种原子任务类型、5 种机器人设置和 2,507 个关节物体。在该数据上预训练的视觉-语言-动作(VLA)模型在仿真基准上表现具竞争力,并实现零样本 sim-to-real 迁移和真实世界关节物体操作的可扩展性能。

正文

View PDF HTML (experimental)

Abstract:The ability to interact with articulated objects is essential for embodied intelligent systems, but collecting large-scale real-world demonstrations for these interactions remains challenging due to the precise contact and constraint-following motions involved. Although simulation provides a promising alternative, existing synthetic data efforts cover limited articulated-object categories, while general-purpose synthesis pipelines lack explicit designs for part-level semantics and articulation constraints, hindering agentic task generation and scalable synthesis of high-quality articulated-manipulation demonstrations. To bridge this gap, we introduce SMART, a scalable system leveraging large-scale Synthesized Manipulation demonstrations for ARTiculated-object manipulation. At its core, we develop SMART-Sim, a simulation platform with articulation-aware design that enables effective task generation and efficient demonstration collection. Building on SMART-Sim, we apply agentic task generation and design a scalable distributed synthesis system, using them to synthesize SMART-Data, comprising over 1M demonstrations across 44 atomic task types, 5 robot setups, and 2,507 articulated objects. The vision-language-action (VLA) model pretrained on SMART-Data shows competitive performance on simulation benchmarks and achieves zero-shot sim-to-real transfer and scalable performance in real-world articulated-object manipulation tasks. This highlights the potential of synthetic demonstrations in providing effective and scalable supervision for improving VLA model performance in contact-rich articulated-object manipulation.
Comments: Technical Report, 31 pages
Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.07652 [cs.RO]
  (or arXiv:2610.07652v1 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2610.07652

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jicong Ao [view email]
[v1] Tue, 6 Oct 2026 02:47:22 UTC (44,530 KB)

来源:arXiv:cs.AI · arxiv.org