跳到正文
arXiv:cs.LG· Hanbin Zhou, Shangzhe Li, Alexander Braverman, Weitong Zhang·· 5 小时前AI 评分33

正则化的可证明收益:对抗模仿学习的快速收敛率

Provable Benefits of Regularization: Fast Rates for Adversarial Imitation Learning

AI 导读

研究提出 Dually Regularized AIL 算法,将 KL 策略正则化与按专家和学习者占用度加权的二次奖励惩罚结合,在有限时域 MDP 与通用函数逼近下证明正则化模仿差距达到 Õ(1/K+1/N) 界。该算法是首个在专家演示和在线交互上同时实现 Õ(1/ε) 样本复杂度的正则化 AIL 方法,即使面对随机专家也成立。

正文

View PDF HTML (experimental)

Abstract:We study adversarial imitation learning (AIL), in which an agent learns to imitate expert demonstrations by optimizing a policy against an adversarial reward that distinguishes expert and learner behavior. Historically, reward regularization and entropy-based policy regularization are key components of empirically successful methods such as GAIL and LS-IQ, yet their finite-sample benefits remain underexplored. We establish fast rates for jointly regularized AIL in finite-horizon Markov decision processes with general function approximation. Our model-free algorithm, Dually Regularized AIL, combines KL policy regularization with a quadratic reward penalty weighted by expert and learner occupancies. With K online episodes and N expert trajectories, we prove a $\widetilde{O}\left(\frac{1}{K}+\frac{1}{N}\right)$ bound on the regularized imitation gap for fixed regularization parameters. Our analysis combines an online mirror descent construction for general convex reward classes to control estimation error from finite expert data and stochastic learner feedback, with a sharp analysis of optimistic KL-regularized policy learning. To the best of our knowledge, Dually Regularized AIL is the first algorithm to simultaneously achieve $\widetilde{O}\left(\frac{1}{\epsilon}\right)$ sample complexity in both expert demonstrations and online interactions for this regularized AIL objective, even with stochastic experts. These results provide a rigorous characterization of the complementary statistical benefits of reward and policy regularization in AIL.
Comments: 33 pages, 1 table
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.35698 [cs.LG]
  (or arXiv:2609.35698v3 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.35698

arXiv-issued DOI via DataCite

Submission history

From: Shangzhe Li [view email]
[v1] Mon, 28 Sep 2026 17:40:11 UTC (50 KB)
[v2] Tue, 29 Sep 2026 16:35:41 UTC (50 KB)
[v3] Fri, 2 Oct 2026 06:35:50 UTC (65 KB)

来源:arXiv:cs.LG · arxiv.org