跳到正文
arXiv:cs.LG· Erik M. Lintunen, Marlos C. Machado·· 3 小时前AI 评分39

Wayfarer:通过发现的选项掌握 Atari 2600 游戏

Mastering Atari 2600 Games with Discovered Options

AI 导读

Wayfarer 是一种通用、领域无关的在线深度强化学习智能体,通过从高维观测中学习 Laplacian 表示来发现选项并用于控制。这些选项能同时改善探索、加速信用分配并有效泛化到未见场景,在最具挑战性的 Atari 2600 游戏中取得单流智能体的 SOTA 表现,在 Montezuma's Revenge 和 Private Eye 等需要长时程探索与策略行为的游戏中提升最大。

正文

View PDF HTML (experimental)

Abstract:Temporal abstractions, often instantiated as options, have long been regarded as a mechanism for accelerating credit assignment, facilitating exploration, and enabling generalisation in reinforcement learning (RL). However, developing general option discovery methods that are effective in large-scale, high-dimensional domains remains a fundamental challenge. Existing option discovery methods are either confined to relatively simple domains, depend on handcrafted or quasi-symbolic representations, or offer little improvement over learning without options. We present Wayfarer, a general, domain-agnostic, online deep RL agent that discovers options through Laplacian representation learning from high-dimensional observations and leverages them for control. We show that the resulting options simultaneously improve exploration, accelerate credit assignment, and generalise effectively to unseen settings, enabling substantially faster learning of complex policies. Wayfarer achieves state-of-the-art performance among single-stream agents on the most challenging Atari 2600 games, with the largest gains in games that require long-horizon exploration and strategic behaviour, such as Montezuma's Revenge and Private Eye.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.03604 [cs.LG]
  (or arXiv:2610.03604v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.03604

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Erik M. Lintunen [view email]
[v1] Fri, 2 Oct 2026 17:05:46 UTC (2,001 KB)

来源:arXiv:cs.LG · arxiv.org