arXiv:cs.AI· Lik Hang Kenny Wong, Yiyao Ma, Xiu-Shen Wei, Zelong Tan, Zhuheng Song, Dongsheng Xie, Kai Chen, Qi Dou·· 4 小时前AI 评分33
VISTA:面向接触密集操作的等变视触觉扩散策略
Equivariant Visual-Tactile Diffusion Policy for Contact-Rich Manipulation
AI 导读
研究者提出 VISTA,一种工作空间级等变视触觉扩散策略,用于数据高效的接触密集模仿学习。该方法将视觉与触觉观测投影为球面 token,通过置换等变球面融合注入触觉接触线索,并用末端执行器朝向旋转融合后的谐波表示,从而预测空间一致的动作。仿真与真实机器人实验显示,VISTA 在数据效率上显著优于强视触觉模仿学习基线,论文已被 CoRL 2026 接收。
正文
Abstract:Imitation learning for contact-rich manipulation requires high-quality expert data that is expensive to obtain. This makes learning a sample-efficient policy a key issue. To address this, we propose VISTA, a workspace-level equivariant visuotactile diffusion policy for data-efficient contact-rich imitation learning. VISTA projects visual and tactile observations into spherical tokens, injects tactile contact cues into visual spherical directions through permutation-equivariant spherical fusion, and rotates the fused harmonic representation using the end-effector orientation. The resulting representation conditions an equivariant diffusion policy to predict spatially consistent actions. Extensive experiments in both simulation and real-world robotic settings show that VISTA substantially improves data efficiency over strong visuotactile imitation learning baselines. Project website: this https URL
| Comments: | 21 pages, 6 figures. Accepted to the 10th Conference on Robot Learning (CoRL 2026) |
| Subjects: | Robotics (cs.RO); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.03333 [cs.RO] |
| (or arXiv:2610.03333v1 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03333 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Lik Hang Kenny Wong [view email]
[v1]
Fri, 2 Oct 2026 14:04:41 UTC (1,818 KB)
来源:arXiv:cs.AI · arxiv.org