arXiv:cs.AI· Haozhuo Zhang, Jingkai Sun, Michele Caprio, Angelo Cangelosi, Jian Tang, Shanghang Zhang, Qiang Zhang, Wei Pan·· 6 小时前AI 评分41
LHM-Humanoid:面向杂乱场景中连续物体搬运的长时程人形运动控制
LHM-Humanoid: Long-Horizon Human Motion Control for Continuous Object Transport in Cluttered Scenes
AI 导读
LHM-Humanoid 用两个目标条件控制器分别完成取物-搬运-放置循环并将终止状态引导至可恢复区域,再通过对抗运动先验正则化并蒸馏为单一策略,实现无重置的连续长时程人形运动。在覆盖四类房间的 350 种杂乱布局上,该方法在已见和未见场景中的成功率和稳定性均大幅优于端到端 RL、分层 RL 及此前的物理仿真人-场景交互方法。
正文
Abstract:Physics-based human motion control can make a simulated character walk, sit, and manipulate objects with high physical realism. Almost always, though, this happens in short, isolated clips that are re-initialized between interactions. We instead aim for continuous, reset-free long-horizon motion: a physically simulated humanoid that repeatedly walks to a displaced object, lifts it with a balanced whole-body posture, carries it past obstacles, and places it at a goal, over and over within a single uninterrupted take. The hard part is not any individual motion but the transitions between them. Without a reset, each cycle must end in a state that both leaves the object just placed undisturbed and lets the next cycle begin, yet every placement leaves the character off-balance in a non-canonical pose where naive end-to-end reinforcement learning fails. Our key idea is to treat this handoff as a two-sided problem of recoverability: the character must disengage from the object it just placed so the prior success is preserved, and settle into a state from which a balanced continuation exists. Instead of engineering a transition by hand, we learn to shape where each cycle ends so that it lands in this recoverable region. We introduce LHM-Humanoid. One goal-conditioned controller completes a fetch--carry--place cycle and, through a learned release-and-retreat behavior, steers its terminal state into this region; a second controller then takes over from the resulting state distribution. Both are regularized by an adversarial motion prior and distilled into a single goal-conditioned policy that runs the whole sequence as one reset-free rollout. Across 350 cluttered layouts spanning four room types, LHM-Humanoid produces far more successful and stable long-horizon motion than end-to-end RL, hierarchical RL, and prior physics-based human-scene-interaction methods, on both seen and unseen scenes.
| Subjects: | Robotics (cs.RO); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2508.16943 [cs.RO] |
| (or arXiv:2508.16943v4 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2508.16943 arXiv-issued DOI via DataCite |
Submission history
From: Haozhuo Zhang [view email]
[v1]
Sat, 23 Aug 2025 08:23:14 UTC (19,322 KB)
[v2]
Thu, 5 Mar 2026 16:38:10 UTC (4,749 KB)
[v3]
Tue, 7 Jul 2026 19:32:44 UTC (4,747 KB)
[v4]
Tue, 6 Oct 2026 12:52:44 UTC (4,757 KB)
来源:arXiv:cs.AI · arxiv.org