跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Tristan Shah, Wooyoung Chung, Volodomyr Makarenko, Juan Wachs, Stas Tiomkin·· 14 小时前AI 评分32

Bellman 遇上 Lyapunov:通过驾驭混沌实现无监督强化学习

Bellman Meets Lyapunov: Unsupervised Reinforcement Learning via Mastering Chaos

AI 导读

研究者提出 F-CIP,一种基于系统动力学本身定义的可控信息生产(CIP)目标的 RL 原生形式,无需人工选择信息变量。该方法与现有 RL 算法兼容,能无监督地发现平衡与维持可控性等基础行为;配合简单的前向速度奖励,还可学得跳跃、奔跑等协调步态。

正文

View PDF HTML (experimental)

Abstract:Reinforcement learning (RL) is a powerful paradigm for training agents, yet its success rests on domain expertise of human engineers who design informative reward signals for every new task. Unsupervised RL aims to reduce this engineering with intrinsic motivation (IM): reward signals that emerge from the agent environment interaction itself. Existing IM objectives, however, involve the selection of information variables, which re-introduces domain expertise the field has sought to eliminate. We introduce Forward CIP (F-CIP), an RL-native formulation of the Controllable Information Production (CIP) objective, which is defined by the system's dynamics alone and requires no such selection. We prove that F-CIP is compatible with RL and demonstrate its effectiveness with existing algorithms. Training agents with F-CIP results in unsupervised discovery of primitive behaviors such as balancing and maintaining controllability, which are essential for more complex robot behaviors. Paired with a simple forward-velocity reward, our method produces coordinated gaits such as hopping and running which otherwise require reward engineering to learn.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.02012 [cs.LG]
  (or arXiv:2610.02012v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02012

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Wooyoung Chung [view email]
[v1] Thu, 1 Oct 2026 16:38:03 UTC (1,760 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org