arXiv:cs.AI(全量分类)· Pedro Robles Dutenhefner, Dikshant Shehmar, Wagner Meira Jr., Marlos C. Machado·· 5 小时前AI 评分36
GITA:为离线目标条件强化学习学习多时间尺度
Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning
AI 导读
针对离线目标条件强化学习(GCRL)中固定时间尺度 k 难以兼顾近距分辨率与长程信号的问题,研究者提出 Generalized Implicit Temporal Abstraction(GITA),让单一价值函数以 k 为条件,并聚合多个 k 值下的优势加权监督来训练策略。
正文
Abstract:Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant states until they fall below the function approximation error, leaving the agent with no signal for ranking states. Temporal abstraction, which treats k environment steps as a single transition, restores this signal at long range, but no single fixed k suits all state-goal distances: large k preserves value differences across long temporal distances while collapsing distinctions between nearby states, and small k does the reverse. We make this trade-off explicit and introduce Generalized Implicit Temporal Abstraction (GITA), which conditions a single value function on k. GITA trains one policy by aggregating advantage-weighted supervision across multiple k values, so scales assigning larger positive advantages to a state-goal pair contribute more strongly to its update. GITA does not need to choose between local resolution and long-range signal; it retains both without committing to a single k. On OGBench, GITA outperforms a broad range of offline GCRL baselines, raising average success rate across all tasks by 25 percentage points (73% relative improvement) over HIQL. It also improves over the strongest fixed-k method, OTA, by 7 percentage points (14% relative).
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.00849 [cs.AI] |
| (or arXiv:2610.00849v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00849 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Pedro Dutenhefner [view email]
[v1]
Thu, 1 Oct 2026 00:06:57 UTC (740 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org