跳到正文
原文
arXiv:cs.AI(全量分类)· Pedro Robles Dutenhefner, Dikshant Shehmar, Wagner Meira Jr., Marlos C. Machado·· 5 小时前AI 评分36

GITA:为离线目标条件强化学习学习多时间尺度

Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning

AI 导读

针对离线目标条件强化学习(GCRL)中固定时间尺度 k 难以兼顾近距分辨率与长程信号的问题,研究者提出 Generalized Implicit Temporal Abstraction(GITA),让单一价值函数以 k 为条件,并聚合多个 k 值下的优势加权监督来训练策略。

正文

View PDF HTML (experimental)

Abstract:Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant states until they fall below the function approximation error, leaving the agent with no signal for ranking states. Temporal abstraction, which treats k environment steps as a single transition, restores this signal at long range, but no single fixed k suits all state-goal distances: large k preserves value differences across long temporal distances while collapsing distinctions between nearby states, and small k does the reverse. We make this trade-off explicit and introduce Generalized Implicit Temporal Abstraction (GITA), which conditions a single value function on k. GITA trains one policy by aggregating advantage-weighted supervision across multiple k values, so scales assigning larger positive advantages to a state-goal pair contribute more strongly to its update. GITA does not need to choose between local resolution and long-range signal; it retains both without committing to a single k. On OGBench, GITA outperforms a broad range of offline GCRL baselines, raising average success rate across all tasks by 25 percentage points (73% relative improvement) over HIQL. It also improves over the strongest fixed-k method, OTA, by 7 percentage points (14% relative).
Subjects: Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
Cite as: arXiv:2610.00849 [cs.AI]
  (or arXiv:2610.00849v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.00849

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Pedro Dutenhefner [view email]
[v1] Thu, 1 Oct 2026 00:06:57 UTC (740 KB)

来源:arXiv:cs.AI(全量分类) · arxiv.org