arXiv:cs.LG(机器学习,全量分类)· Jade Choghari, Pepijn Kooijmans, Mansi Agarwal, Yusuf Umut Ciftci, Aseem Doriwala, Catherine Weaver, Mouli Sivapurapu, Kai Yang, Thomas Wolf, Jackson Lee, Pragna Mannam·· 5 小时前AI 评分52
Hugging Face 团队发布 FineART 双臂操作数据集与 FineART-VLA 模型
FineART: Fine-Grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation
AI 导读
Hugging Face 团队发布 FineART 双臂操作数据集,含 40,543 条轨迹(1,718 小时)、151 个任务共 533,913 个子任务标注,并推出能预测下一子任务来引导动作的 FineART-VLA 模型。
正文
Authors:Jade Choghari, Pepijn Kooijmans, Mansi Agarwal, Yusuf Umut Ciftci, Aseem Doriwala, Catherine Weaver, Mouli Sivapurapu, Kai Yang, Thomas Wolf, Jackson Lee, Pragna Mannam
Abstract:Robots operating in real-world environments must often execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets struggle to support this capability: although single-arm datasets reach hundreds of thousands of trajectories, they typically provide only one high-level instruction per episode, while existing bimanual datasets with subtask labels annotate only part of their recorded hours. We present FineART, a densely annotated bimanual manipulation dataset comprising 40,543 episodes (1,718 hours) and 533,913 subtasks across 151 tasks. We also introduce FineART-VLA, a vision-language-action policy that predicts its own next subtask to guide its actions. Mid-training on FineART's subtask annotations raises FineART-VLA's success at following spatial instructions from 32.0% to 100.0%. With step-by-step human subtask guidance, it also raises success on unseen long-horizon tasks from 16.0% to 76.0%. Furthermore, after minimal fine-tuning on a new robot, the policy requires only one-tenth of the data needed by baselines without this mid-training and generalizes zero-shot to tasks unseen on the new hardware. We open-source the full dataset, model weights, and training code.
| Comments: | 26 pages. Code and model weights will be integrated into Hugging Face LeRobot this https URL |
| Subjects: | Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.36416 [cs.RO] |
| (or arXiv:2609.36416v2 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2609.36416 arXiv-issued DOI via DataCite |
Submission history
From: Jade Choghari [view email]
[v1]
Tue, 29 Sep 2026 00:13:29 UTC (5,967 KB)
[v2]
Wed, 30 Sep 2026 23:07:44 UTC (5,967 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org