跳到正文
arXiv:cs.AI· Kush Hari, Justin Kerr, Nidhya Shivakumar, Samarth Mahapatra, Carmelo Sferrazza, Jiahui Lei, Jitendra Malik, C. Karen Liu, Ken Goldberg, Angjoo Kanazawa·· 4 小时前AI 评分48

EyeRobot 2.0:无需腕部相机,用主动注视实现精准操作

EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras

AI 导读

EyeRobot 2.0 提出用主动注视(Active Visual Fixation)框架,仅靠单个立体相机完成精细双手操作,通过转动两个视点将注视中心对准 3D 注视点,并以中央凹方式分配更多视觉 token。

正文

View PDF HTML (experimental)

Abstract:Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it. The resulting images are processed foveally by allocating more visual tokens to the image centers, focusing computation on task-relevant features. Such Active Visual Fixation (AVF) requires carefully coordinated gaze during task execution, which we accomplish hierarchically by first training a low-level gaze servoing policy conditioned on a goal object, then training a target selector which emits fixation goals based on task progress. Both modules are trained with RL on real-world data: the first is trained with a dense geometric reward and the second co-trains with the BC gripper policy which allows it to discover fixation sequences that can resemble a human's fixation sequence while performing the task. EyeRobot 2.0 further takes advantage of fixation by canonicalizing gripper information into a fixation-relative SE(3) frame, which compacts the size of the action distribution to learn. We collect teleoperation data for 7 real-world and 6 simulated tasks, and conduct over 1000 physical and 1800 simulated robot trials comparing EyeRobot 2.0 against passive stereo and ego + wrist camera policies trained on the same data. Removing wrist cameras is costly for standard policies: with only passive stereo, real-world success drops from 52% to 27%. EyeRobot 2.0 closes this gap with only stereo, outperforming passive stereo by 40% in real and 20% in sim. It matches ego + wrist policies when their wrist views are clear (69% vs. 64%), and more than doubles their success when grasped objects occlude the wrist cameras (48% vs. 22%)
Comments: Project Page: this https URL
Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.03710 [cs.RO]
  (or arXiv:2610.03710v1 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2610.03710

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Justin Kerr [view email]
[v1] Fri, 2 Oct 2026 17:57:51 UTC (19,700 KB)

来源:arXiv:cs.AI · arxiv.org