arXiv:cs.LG· Jiachen Zhao, Antonia Januszewicz, Taeho Jung·· 4 小时前AI 评分43
提示词级差分隐私下的奖励驱动学习:RLVR 训练的首个差分隐私保证
Reward-Driven Learning under Prompt-Level Differential Privacy
AI 导读
研究提出首个针对 RLVR(可验证奖励强化学习)训练的差分隐私保证,采用 prompt 级差分隐私,将同一 prompt 的所有回复梯度聚合、对 prompt 贡献裁剪一次并加高斯噪声。
正文
Abstract:Reinforcement learning with verifiable rewards (RLVR) trains a language model on problems that may themselves be confidential, and the trained model can reveal which problems it saw. We study RLVR under prompt-level differential privacy: the released weights must be ({\epsilon},{\delta})-differentially private with respect to the presence of any one training problem. Taking the group of responses to one prompt as the privacy record, our method aggregates their gradients, clips the prompt's contribution once, adds Gaussian noise, and composes the privacy loss across updates, so the budget depends on neither the number of responses per prompt nor the clipping norm; to our knowledge this is the first differential privacy guarantee for RLVR training. We train Qwen2.5-1.5B-Instruct with LoRA at a per-run budget of {\epsilon}=8 and compare, on the same prompts and at the same budget, a control that removes only the reward signal and two private supervised fine-tuning recipes. The reward signal improves accuracy over the control by 2.65 points on MATH and 3.24 on GSM8K, in every seed; the improvement survives a format-robust scorer, at 1.3 points on MATH, and is not explained by response length. At the same budget the private model outperforms both supervised recipes on MATH and GSM8K by 2.3 to 3.8 points, retains 85--90% of the gain of non-private GRPO on these tasks, and on MATH the noise of an eightfold tighter budget costs at most 1.2 points. The reward effect also carries to CommonsenseQA, an exploratory non-mathematical task. Verifier feedback thus remains a usable learning signal under prompt-level privacy.
| Subjects: | Machine Learning (cs.LG); Cryptography and Security (cs.CR) |
| Cite as: | arXiv:2610.07212 [cs.LG] |
| (or arXiv:2610.07212v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07212 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jiachen Zhao [view email]
[v1]
Mon, 5 Oct 2026 18:26:21 UTC (179 KB)
来源:arXiv:cs.LG · arxiv.org