arXiv:cs.LG· Giacomo Spigler·· 5 小时前AI 评分46
TAVIS:面向模仿学习的自我中心主动视觉与预期注视基准
TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning
AI 导读
TAVIS 是用于主动视觉模仿学习的评测基础设施,包含 TAVIS-Head(5 项任务,pan/tilt 颈部全局搜索)和 TAVIS-Hands(3 项任务,腕部相机局部遮挡)两套任务,基于 IsaacLab,覆盖 GR1T2 与 Reachy2 两种人形躯干。
正文
Abstract:Active vision -- where a policy controls its own gaze during manipulation -- has emerged as a key capability for imitation learning, with multiple independent systems demonstrating its benefits in the past year. Yet there is no shared benchmark to compare approaches or quantify what active vision contributes, on which task types, and under what conditions. We introduce TAVIS, evaluation infrastructure for active-vision imitation learning, with two complementary task suites -- TAVIS-Head (5 tasks, global search via pan/tilt necks) and TAVIS-Hands (3 tasks, local occlusion via wrist cameras) -- on two humanoid torso embodiments (GR1T2, Reachy2), built on IsaacLab. TAVIS provides three evaluation primitives: a paired headcam-vs-fixedcam protocol on identical demonstrations; GALT (Gaze-Action Lead Time), a novel metric grounded in cognitive science and HRI that quantifies anticipatory gaze in learned policies; and procedural ID/OOD splits. Baseline experiments with Diffusion Policy and $\pi_0$ reveal that (i) active vision generally helps, consistently across training runs, but benefits are task-conditional rather than uniform; (ii) multi-task policies degrade sharply under controlled distribution shifts on both suites; and (iii) imitation alone yields anticipatory gaze that lands on the task-relevant object, with lead times comparable to those of the human demonstrations, although head motion is less smooth than in the demonstrations, an effect of action chunking that success rate does not reveal. Code and evaluation scripts are released at this https URL demonstrations (LeRobot v3.0; ~2200 episodes) and trained baselines at this https URL.
| Comments: | 29 pages, 6 figures, 14 tables. v2: revised and extended (inter-seed variability, wrist-camera ablation, GALT validation and sensitivity analyses, gaze behaviour statistics). Project page: this https URL |
| Subjects: | Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2605.07943 [cs.RO] |
| (or arXiv:2605.07943v2 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2605.07943 arXiv-issued DOI via DataCite |
Submission history
From: Giacomo Spigler [view email]
[v1]
Fri, 8 May 2026 16:11:13 UTC (7,612 KB)
[v2]
Tue, 6 Oct 2026 18:52:01 UTC (8,556 KB)
来源:arXiv:cs.LG · arxiv.org