arXiv:cs.LG(机器学习,全量分类)· Yue YU, Bowen Zuo, David Crandall, Yinglun Zhu, Dongruo Zhou·· 14 小时前AI 评分34
StepRS-GRPO:面向扩散语言模型的分步风险敏感 GRPO 方法
Exploring More, Reasoning Better: Stepwise Risk-Sensitive GRPO for Diffusion Language Models
AI 导读
研究者提出 StepRS-GRPO,一种面向扩散大语言模型(dLLM)的分步风险敏感 GRPO 方法,通过在去噪状态间调整组优势变换的风险系数来改进强化学习。在多个 dLLM 骨干和数学推理基准上,该方法相比居中 GRPO 同时提升了 pass@1 准确率和 pass@k 覆盖率,并增加了答案多样性。
正文
Abstract:Diffusion large language models (dLLMs) generate text by denoising a sequence or successive blocks, allowing several tokens to be revealed in parallel. Reinforcement learning with verifiable rewards (RLVR) reuses terminal feedback across these decisions, even as their conditioning context changes. We propose stepwise risk-sensitive GRPO (StepRS-GRPO), which varies the risk coefficient of the group-advantage transformation across denoising states while retaining the underlying trainer. For binary rewards, we show that this transformation is exactly a prompt- and state-dependent rescaling of centered outcome advantages. A capability-based calibration suggests a coefficient scale, while endpoint and interpolation ablations guide schedule selection. Across multiple dLLM backbones and mathematical reasoning benchmarks, StepRS-GRPO improves both pass@1 accuracy and pass@k coverage over centered GRPO, while increasing answer diversity. In our ablation studies, mass-matched controls support the contributions of state allocation and schedule direction, and the gains persist after matching the root mean square (RMS) of the advantages to that of centered GRPO. Reasoning-trace diagnostics further show that the diversity gains from StepRS-GRPO extend beyond final-answer strings.
| Comments: | 40 pages, 12 figures, 2 tables. The first two authors contributed equally |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.00661 [cs.LG] |
| (or arXiv:2610.00661v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00661 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yue Yu [view email]
[v1]
Wed, 30 Sep 2026 19:59:48 UTC (665 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org