跳到正文
arXiv:cs.LG· SungJae Ahn, Jeong Woon Lee, Kyoleen Kwak, Hyoseok Hwang·· 4 小时前AI 评分34

重新审视深度强化学习中的时序正则化以实现平滑控制

Revisiting Temporal Regularization for Smooth Control in Deep Reinforcement Learning

AI 导读

研究证明时序惩罚项能约束共享下一状态的当前状态间的预期动作差异,从而提供空间平滑效果,据此提出仅用时序平滑进行动作条件化(CATS)方法,将时序惩罚与线性递增结合。在仿真和真实环境实验中,CATS 显著减少动作振荡且不降低任务性能,计算开销很小。

正文

View PDF HTML (experimental)

Abstract:Deep Reinforcement Learning policies can produce nonsmooth action oscillations that hinder deployment on physical robots. Existing architectural and penalty-based approaches seek spatial smoothness by directly reducing sensitivity to changes in state inputs, but their broad constraints can degrade task performance as stronger smoothing is pursued. Temporal regularization instead constrains action differences along observed transitions, but has been considered unable to provide the spatial smoothness needed under observation noise. We revisit this assumption by proving that the temporal penalty bounds the expected action differences between current states sharing a next state, revealing a spatial effect that empirically extends to spatial smoothness. Building on this finding, we propose Conditioning for Action using only Temporal Smoothness (CATS), which combines a temporal penalty with linear ramp-up. We highlight temporal regularization's ability to provide spatial smoothness while better preserving task performance than explicit spatial regularization. Through linear ramp-up, CATS allows the policy to learn rewarding behavior before progressively smoothing its actions, improving return preservation and both temporal and spatial smoothness. Experiments in both simulation and the real world show that CATS substantially reduces action oscillation without degrading task performance, with little computational overhead.
Comments: Submitted to ICRA 2027
Subjects: Machine Learning (cs.LG); Robotics (cs.RO)
Cite as: arXiv:2610.07910 [cs.LG]
  (or arXiv:2610.07910v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07910

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: SungJae Ahn [view email]
[v1] Tue, 6 Oct 2026 07:53:24 UTC (631 KB)

来源:arXiv:cs.LG · arxiv.org