arXiv:cs.AI· Lihan Zha, Shresth Grover, Tenny Yin, Samuel M. Bateman, Hengkai Pan, Mengchao Zhang, Aykut Onol, Allen Z. Ren, Dhruv Shah, Anirudha Majumdar·· 6 小时前AI 评分42
EgoLAP:通过语言-动作推理从第一视角人类数据中学习
EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning
AI 导读
EgoLAP 是一个 VLA 预训练框架,通过共享的基于语言的动作链式推理,从人类和机器人轨迹中联合学习。它将运动意图表达为结构化、时间抽象的语言动作,并配以基于场景几何、物理和物体可供性的运动级推理。在真实世界与仿真实验中,EgoLAP 达到 80.1% 的平均真实任务进度,比替代动作表示提升 2.3 倍,且运动级推理优于结合子任务、物体框和视觉轨迹的复合推理格式。
正文
Abstract:Egocentric human data offer a path to scaling robot learning beyond costly robot demonstrations, yet the embodiment gap makes raw human trajectories a poor supervisory target for control. Our key insight is that, although low-level actions are embodiment-specific, their underlying motion intent can capture task-relevant structure that transfers across humans and robots. We introduce EgoLAP, a VLA pre-training framework that jointly learns from human and robot trajectories through a shared language-based action chain-of-thought. EgoLAP expresses motion intent as structured, temporally abstracted language actions and pairs them with motion-level reasoning grounded in scene geometry, physics, and object affordances. Across extensive real-world and simulated experiments, EgoLAP transfers human experience to robot control more effectively than alternative action representations and reaches 80.1% mean real-world task progress, a 2.3x performance gain over alternative action representations. Motion-level reasoning also outperforms a composite reasoning format that combines subtask, object-box, and visual-trace reasoning.
| Comments: | Project website: this https URL |
| Subjects: | Robotics (cs.RO); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08726 [cs.RO] |
| (or arXiv:2610.08726v1 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08726 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Lihan Zha [view email]
[v1]
Tue, 6 Oct 2026 17:30:33 UTC (13,438 KB)
来源:arXiv:cs.AI · arxiv.org