跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Daisuke Yamada, Travis Pence, Vikas Singh·· 15 小时前AI 评分30

ReQRL:从有限时间可达性学习目标可达的拟度量几何

Learning Goal-Reaching Quasimetric Geometry From Finite-Time Reachability

AI 导读

研究者提出 ReQRL,通过有限时间可达性约束 critic 的价值梯度,从数据中解耦动力学可达性与边界几何,用于目标条件强化学习中的拟度量学习。在 OGBench 上,该方法优于或媲美现有拟度量方法及其他离线 GCRL 方法。

正文

View PDF HTML (experimental)

Abstract:In goal-conditioned reinforcement learning (GCRL), quasimetric learning models goal-reaching costs as quasimetric distances, connecting local constraints to global value geometry. Its local constraints, however, should reflect the direction- dependent effects of control composition over a finite horizon together with environmental feasibility. We propose ReQRL, which constrains the critic's value gradients through finite-horizon reachability. Drawing on state-constrained optimal control, we decouple dynamical reachability from boundary geometry, estimating both from data. On OGBench, our method outperforms or rivals existing quasimetric approaches and other offline GCRL methods.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.00778 [cs.LG]
  (or arXiv:2610.00778v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00778

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Daisuke Yamada [view email]
[v1] Wed, 30 Sep 2026 22:16:14 UTC (1,355 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org