跳到正文
arXiv:cs.AI· Clarisse Wibault, Antoine Gorceix, Antonio L\'eon Villares, Alexey Zakharov, Evangelos Chatzaroulas, Michael Matthews, Eduardo Pignatelli, Jakob Foerster·· 3 小时前

Q-Shaped Options:面向分层强化学习的动作抽象方法

Q-Shaped Options for Hierarchical Reinforcement Learning

AI 导读

研究者提出 Q-Shaped Options(QSO),通过由各层级 Q 函数共同塑造的共享编码器学习动作抽象,以同时满足动作抽象的三项要求。QSO 在分层架构中让低层 Q 函数将 option 作为目标、高层 Q 函数将其作为动作,从而兼顾最优控制所需的区分与不必要的区分去除。

正文

View PDF

Abstract:Learning to tackle long-horizon, goal-conditioned tasks requires an agent to reason over extended timescales and act across a broad range of states. In principle, Hierarchical Reinforcement Learning (HRL) addresses both challenges through the interaction between action (temporal) and state (spatial) abstraction. First, using an action abstraction to represent temporally extended behaviour as options reduces the effective decision horizon. Second, enabling different state abstractions at each level of the decision process permits greater data aggregation for learning. However, realising these two benefits of a hierarchical policy depends on learning an appropriate action abstraction. Current HRL algorithms fail in one of two ways. Some discard distinctions between options needed for optimal control, undermining hierarchy altogether. Others retain unnecessary distinctions, preserving horizon reduction, but forfeiting coarser state abstraction. In this work, we characterise three desiderata for an action abstraction. We introduce Q-Shaped Options (QSO) to address all three. QSO builds on an architecture with distinct state-value functions, Q functions and policies at each level of the hierarchy. It learns the action abstraction between consecutive levels as a shared encoder shaped by their respective Q functions. The low-level Q function uses the option as a goal, encouraging the abstraction to retain distinctions necessary for optimal control. The high-level Q function uses it as an action, encouraging unnecessary distinctions to be discarded. Across offline goal-conditioned locomotion and manipulation environments, QSO learns semantically meaningful option spaces and outperforms baselines, achieving non-zero performance in tasks where all other evaluated algorithms fail.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.12135 [cs.AI]
  (or arXiv:2610.12135v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.12135

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Clarisse Wibault [view email]
[v1] Thu, 8 Oct 2026 15:23:39 UTC (6,110 KB)

来源:arXiv:cs.AI · arxiv.org