Apple Machine Learning Research(RSS)·· 15 天前AI 评分40
DACA-GRPO:面向扩散语言模型强化学习的去噪感知信用分配方法
DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models
AI 导读
Apple 研究者提出 DACA-GRPO,一种可插拔的 GRPO 训练增强方法,用于扩散语言模型强化学习。它通过 Denoising Progress Scores 提取逐 token 重要性权重(无额外前向开销),并用 Stratified Masking Likelihood 降低 mean-field 似然偏差。
来源:Apple Machine Learning Research(RSS) · machinelearning.apple.com