跳到正文
arXiv:cs.LG· Ruimin He, Shaowei Lin·· 4 小时前AI 评分28

动作驱动过程:统一随机过程与强化学习的连续时间控制框架

Action-Driven Processes for Continuous-Time Control

AI 导读

论文提出“动作驱动过程”框架,将随机过程与强化学习统一起来,并应用于脉冲神经网络。基于 control-as-inference 思路,作者证明:在恰当定义的动作驱动过程上,最小化策略驱动的真实分布与奖励驱动的模型分布之间的 KL 散度,等价于最大熵强化学习。

正文

View PDF HTML (experimental)

Abstract:At the heart of reinforcement learning are actions -- decisions made in response to observations of the environment. Actions are equally fundamental in the modeling of stochastic processes, as they trigger discontinuous state transitions and enable the flow of information through large, complex systems. In this paper, we unify the perspectives of stochastic processes and reinforcement learning through action-driven processes, and illustrate their application to spiking neural networks. Leveraging ideas from control-as-inference, we show that minimizing the Kullback-Leibler divergence between a policy-driven true distribution and a reward-driven model distribution for a suitably defined action-driven process is equivalent to maximum entropy reinforcement learning.
Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG)
Cite as: arXiv:2510.26672 [stat.ML]
  (or arXiv:2510.26672v3 [stat.ML] for this version)
  https://doi.org/10.48550/arXiv.2510.26672

arXiv-issued DOI via DataCite

Submission history

From: Shaowei Lin [view email]
[v1] Thu, 30 Oct 2025 16:42:09 UTC (24 KB)
[v2] Sat, 26 Sep 2026 00:34:44 UTC (24 KB)
[v3] Mon, 5 Oct 2026 19:09:24 UTC (24 KB)

来源:arXiv:cs.LG · arxiv.org