arXiv:cs.CL· Fei Ding·· 3 小时前
为何 KL 正则化在组策略优化中失效:ZCPO 提出条件 KL 校准方案
When KL Regularization Misfires in Group Policy Optimization
AI 导读
研究分析了参考策略 KL 正则化与奖励交互中的七种失效模式,包括奖励裁剪后的残余 KL 更新、梯度抵消、组内奖励相同时的 KL 更新、KL 随响应长度增长及相对贡献失衡、KL 集中于少量 token,以及 k1 融入奖励时的采样噪声。
正文
Abstract:Why does removing reference-policy KL regularization sometimes improve group policy optimization? This motivates studying how reference-policy information should enter group-relative updates. We analyze seven potential failure modes in the interactions between KL and rewards: residual KL updates after reward clipping, after gradient cancellation, and in groups with identical rewards; KL growth with response length and an imbalance in its relative contribution; KL concentration on a small number of tokens; and sampling noise when k1 is incorporated into rewards. We propose Zero-Sum Calibrated Policy Optimization (ZCPO), which uses relative drift measured by conditional KL to calibrate within-group reward coefficients and integrates them into the base surrogate. Mathematical reasoning experiments and ablations support this design's effectiveness in our settings.
| Comments: | 27 pages |
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL) |
| MSC classes: | 68T05, 68T50 |
| ACM classes: | I.2.6; I.2.7 |
| Cite as: | arXiv:2610.12161 [cs.LG] |
| (or arXiv:2610.12161v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12161 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ding Fei [view email]
[v1]
Thu, 8 Oct 2026 15:39:09 UTC (2,295 KB)
来源:arXiv:cs.CL · arxiv.org