跳到正文
arXiv:cs.AI· Kevin Yandoka Denamgana\"i, Daniel Hernandez, Ozan Vardal, Sondess Missaoui, James Alfred Walker·· 4 小时前AI 评分33

ETHER:用涌现通信对齐 HER 的文本后见经验回放

ETHER: Aligning Emergent Communication for Hindsight Experience Replay

AI 导读

ETHER(Emergent Textual Hindsight Experience Replay)通过涌现通信学习目标重标注与谓词函数,解决指令跟随任务中 HER 依赖预设函数的问题。

正文

View PDF HTML (experimental)

Abstract:Hindsight Experience Replay (HER) enhances sample efficiency in goal-conditioned reinforcement learning (RL) by relabelling failed trajectories with goals that were actually achieved. However, HER implicitly assumes access to a goal relabelling function and a predicate function that determines whether a goal has been satisfied. These assumptions break down in instruction-following tasks, where goals are expressed in natural language and differ from the state space. We formalize this as the Hindsight Reinforcement Learning problem, which shows the need to jointly learn these functions alongside the RL policy. To address it, we propose ETHER (Emergent Textual Hindsight Experience Replay), an agent that leverages Emergent Communication to learn the goal-relabelling and predicate functions. ETHER uses a referential game (RG) to train a speaker and a listener to develop a grounded, artificial language describing environment states. It partially aligns this emergent language with instruction language using co-occurrence patterns between task instructions and RL observations. We prove that the relabelling and predicate functions that ETHER derives from the RG avoid the degenerate solutions of the Hindsight RL problem, namely trivial predicates and collapsed relabelling functions. Experiments on BabyAI's PickupDist task show that ETHER's learned RG speaker and listener can function as the goal relabelling and predicate functions of HER, improving sample efficiency despite imperfect language alignment. Our work bridges Emergent Communication and goal-conditioned RL, opening the door to wider applications of HER.
Comments: work in progress
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
Cite as: arXiv:2307.15494 [cs.CL]
  (or arXiv:2307.15494v4 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2307.15494

arXiv-issued DOI via DataCite

Submission history

From: Kevin Denamganaï [view email]
[v1] Fri, 28 Jul 2023 11:42:31 UTC (5,068 KB)
[v2] Sun, 17 Dec 2023 10:30:11 UTC (5,778 KB)
[v3] Wed, 30 Sep 2026 13:49:33 UTC (5,017 KB)
[v4] Fri, 2 Oct 2026 16:59:53 UTC (5,021 KB)

来源:arXiv:cs.AI · arxiv.org