跳到正文
arXiv:cs.AI· Haozhuo Zhang, Jingkai Sun, Michele Caprio, Angelo Cangelosi, Jian Tang, Shanghang Zhang, Qiang Zhang, Wei Pan·· 6 小时前AI 评分41

LHM-Humanoid:面向杂乱场景中连续物体搬运的长时程人形运动控制

LHM-Humanoid: Long-Horizon Human Motion Control for Continuous Object Transport in Cluttered Scenes

AI 导读

LHM-Humanoid 用两个目标条件控制器分别完成取物-搬运-放置循环并将终止状态引导至可恢复区域,再通过对抗运动先验正则化并蒸馏为单一策略,实现无重置的连续长时程人形运动。在覆盖四类房间的 350 种杂乱布局上,该方法在已见和未见场景中的成功率和稳定性均大幅优于端到端 RL、分层 RL 及此前的物理仿真人-场景交互方法。

正文

View PDF HTML (experimental)

Abstract:Physics-based human motion control can make a simulated character walk, sit, and manipulate objects with high physical realism. Almost always, though, this happens in short, isolated clips that are re-initialized between interactions. We instead aim for continuous, reset-free long-horizon motion: a physically simulated humanoid that repeatedly walks to a displaced object, lifts it with a balanced whole-body posture, carries it past obstacles, and places it at a goal, over and over within a single uninterrupted take. The hard part is not any individual motion but the transitions between them. Without a reset, each cycle must end in a state that both leaves the object just placed undisturbed and lets the next cycle begin, yet every placement leaves the character off-balance in a non-canonical pose where naive end-to-end reinforcement learning fails. Our key idea is to treat this handoff as a two-sided problem of recoverability: the character must disengage from the object it just placed so the prior success is preserved, and settle into a state from which a balanced continuation exists. Instead of engineering a transition by hand, we learn to shape where each cycle ends so that it lands in this recoverable region. We introduce LHM-Humanoid. One goal-conditioned controller completes a fetch--carry--place cycle and, through a learned release-and-retreat behavior, steers its terminal state into this region; a second controller then takes over from the resulting state distribution. Both are regularized by an adversarial motion prior and distilled into a single goal-conditioned policy that runs the whole sequence as one reset-free rollout. Across 350 cluttered layouts spanning four room types, LHM-Humanoid produces far more successful and stable long-horizon motion than end-to-end RL, hierarchical RL, and prior physics-based human-scene-interaction methods, on both seen and unseen scenes.
Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI)
Cite as: arXiv:2508.16943 [cs.RO]
  (or arXiv:2508.16943v4 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2508.16943

arXiv-issued DOI via DataCite

Submission history

From: Haozhuo Zhang [view email]
[v1] Sat, 23 Aug 2025 08:23:14 UTC (19,322 KB)
[v2] Thu, 5 Mar 2026 16:38:10 UTC (4,749 KB)
[v3] Tue, 7 Jul 2026 19:32:44 UTC (4,747 KB)
[v4] Tue, 6 Oct 2026 12:52:44 UTC (4,757 KB)

来源:arXiv:cs.AI · arxiv.org