跳到正文
arXiv:cs.CL· Fei Ding·· 3 小时前

为何 KL 正则化在组策略优化中失效:ZCPO 提出条件 KL 校准方案

When KL Regularization Misfires in Group Policy Optimization

AI 导读

研究分析了参考策略 KL 正则化与奖励交互中的七种失效模式,包括奖励裁剪后的残余 KL 更新、梯度抵消、组内奖励相同时的 KL 更新、KL 随响应长度增长及相对贡献失衡、KL 集中于少量 token,以及 k1 融入奖励时的采样噪声。

正文

View PDF HTML (experimental)

Abstract:Why does removing reference-policy KL regularization sometimes improve group policy optimization? This motivates studying how reference-policy information should enter group-relative updates. We analyze seven potential failure modes in the interactions between KL and rewards: residual KL updates after reward clipping, after gradient cancellation, and in groups with identical rewards; KL growth with response length and an imbalance in its relative contribution; KL concentration on a small number of tokens; and sampling noise when k1 is incorporated into rewards. We propose Zero-Sum Calibrated Policy Optimization (ZCPO), which uses relative drift measured by conditional KL to calibrate within-group reward coefficients and integrates them into the base surrogate. Mathematical reasoning experiments and ablations support this design's effectiveness in our settings.
Comments: 27 pages
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL)
MSC classes: 68T05, 68T50
ACM classes: I.2.6; I.2.7
Cite as: arXiv:2610.12161 [cs.LG]
  (or arXiv:2610.12161v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.12161

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ding Fei [view email]
[v1] Thu, 8 Oct 2026 15:39:09 UTC (2,295 KB)

来源:arXiv:cs.CL · arxiv.org