跳到正文
HuggingFace Daily Papers·· 1 天前AI 评分40

Attacca:面向长时程具身智能体的状态连续性目标导向控制

Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents

AI 导读

针对长时程具身任务中目标常不在视野内、场景与位姿随任务递变的问题,Attacca 提出用与执行环境解耦的目标图像训练视觉目标条件策略,并通过目标掩码预测头学习当前视角的密集定位、用行为阶段条件区分 Search、Approach、Interact 三阶段。

正文

View PDF HTML (experimental)

Abstract:A central capability of embodied agents is to accomplish complex objectives through sequences of interdependent tasks. Yet existing visual goal-conditioned policies underlying these agents are typically evaluated on isolated interactions where the target is already visible, and thus do not capture the conditions that arise during continuous long-horizon task execution. In such settings, each task begins from the state left by the previous one: the agent may end at a different position and orientation, the world may have been modified, and the next interaction target may lie outside the current field of view. As a result, agents relying on such policies may struggle to proceed to the next task when they cannot ground their target in the current observation. To address this challenge, we propose Attacca, a new approach that trains visual goal-conditioned policies on complete search-to-interact trajectories using goal images decoupled from the execution environment. Attacca uses context-decoupled goal sampling to pair each demonstration with a class-compatible masked goal image from another world, removing direct scene and pose correspondence. It learns dense current-view grounding through a target-mask prediction head, providing auxiliary supervision beyond action imitation. We further introduce behavioral-phase conditioning that teaches the policy to distinguish Search, Approach, and Interact stages and adapt its control as execution progresses. We evaluate Attacca on multiple short- and long-horizon embodied tasks in Minecraft. Our method achieves 39.0-47.5% clean success, improving over the strongest baseline by 1.7-2.4x. On long-horizon tasks, it attains 54%, 30%, and 28% completion, yielding up to a 7x improvement.
Comments: Project page: this https URL
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.07785 [cs.AI]
  (or arXiv:2610.07785v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.07785

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Gyusik Seo [view email]
[v1] Tue, 6 Oct 2026 05:29:18 UTC (22,969 KB)

来源:HuggingFace Daily Papers · arxiv.org