跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Amr Mousa, Rifny Rachman, Neil Karavis, Michele Caprio, Richard Allmendinger·· 9 小时前AI 评分36

PROMO:面向四足机器人的偏好条件多目标强化学习

PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots

AI 导读

PROMO 将四足机器人运动中的目标权衡变为运行时输入,用单一策略替代固定标量奖励。在仿真中采样 100 个偏好,67 个行为处于 Pareto 非支配状态,偏好与目标平均相关性 0.843。同一策略零样本迁移至 Unitree Go2,仅改变偏好即可将特定能耗降低最多 30.4%、位置误差降低 38.7%、机体姿态峰值偏差降低 59.0%。

正文

View PDF HTML (experimental)

Abstract:Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy. PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion priors fixed, thereby separating operator intent from reward shaping terms required for viable gait generation. Compared with fixed-objective controllers, multi-objective baselines, and independently trained specialists, PROMO achieves objective specialization and robustness from a single deployable policy. Across 100 sampled preferences in simulation, 67 behaviors are non-dominated under exact Pareto dominance, with a mean preference-objective correlation of 0.843, demonstrating broad Pareto coverage and predictable preference response. The same policy transfers zero-shot to a Unitree Go2, where preference changes alone reduce specific energy by up to 30.4%, position error by 38.7%, and peak body-attitude deviation by 59.0% relative to the balanced preference. These results establish preference-conditioned multi-objective RL as a practical runtime interface for adaptive legged locomotion, extending its role beyond offline Pareto-set construction. Open-source code and videos are available at this https URL.
Comments: Submitted to IEEE Transactions on Robotics. Project website, code, and videos: this https URL
Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG); Systems and Control (eess.SY)
Cite as: arXiv:2610.01260 [cs.RO]
  (or arXiv:2610.01260v1 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2610.01260

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Amr Mousa [view email]
[v1] Thu, 1 Oct 2026 07:52:33 UTC (9,630 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org