arXiv:cs.LG(机器学习,全量分类)· Amr Mousa, Rifny Rachman, Neil Karavis, Michele Caprio, Richard Allmendinger·· 9 小时前AI 评分36
PROMO:面向四足机器人的偏好条件多目标强化学习
PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots
AI 导读
PROMO 将四足机器人运动中的目标权衡变为运行时输入,用单一策略替代固定标量奖励。在仿真中采样 100 个偏好,67 个行为处于 Pareto 非支配状态,偏好与目标平均相关性 0.843。同一策略零样本迁移至 Unitree Go2,仅改变偏好即可将特定能耗降低最多 30.4%、位置误差降低 38.7%、机体姿态峰值偏差降低 59.0%。
正文
Abstract:Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy. PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion priors fixed, thereby separating operator intent from reward shaping terms required for viable gait generation. Compared with fixed-objective controllers, multi-objective baselines, and independently trained specialists, PROMO achieves objective specialization and robustness from a single deployable policy. Across 100 sampled preferences in simulation, 67 behaviors are non-dominated under exact Pareto dominance, with a mean preference-objective correlation of 0.843, demonstrating broad Pareto coverage and predictable preference response. The same policy transfers zero-shot to a Unitree Go2, where preference changes alone reduce specific energy by up to 30.4%, position error by 38.7%, and peak body-attitude deviation by 59.0% relative to the balanced preference. These results establish preference-conditioned multi-objective RL as a practical runtime interface for adaptive legged locomotion, extending its role beyond offline Pareto-set construction. Open-source code and videos are available at this https URL.
| Comments: | Submitted to IEEE Transactions on Robotics. Project website, code, and videos: this https URL |
| Subjects: | Robotics (cs.RO); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG); Systems and Control (eess.SY) |
| Cite as: | arXiv:2610.01260 [cs.RO] |
| (or arXiv:2610.01260v1 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01260 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Amr Mousa [view email]
[v1]
Thu, 1 Oct 2026 07:52:33 UTC (9,630 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org