跳到正文
arXiv:cs.CL· Blake Bullwinkel, Eugenia Kim, Amanda Minnich, Mark Russinovich·· 4 小时前

用 GRPO 实现语言模型的自适应红队攻防

Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO

AI 导读

研究探索用 GRPO 训练语言模型的攻击方与防御方,通过多个基于 LLM 评判的奖励通道并用 GDPO 计算优势以避免单一通道主导。方法采用从攻击方单轮、多轮训练到攻防交替协同训练的课程,产出攻击效果强且可迁移,协同训练的防御方在保持通用效用的同时达到有竞争力的安全性。消融实验识别了影响安全与效用平衡的关键组件,并发现 GRPO 会在训练中令攻击方多样性坍缩。

正文

View PDF HTML (experimental)

Abstract:Language model safety must continually adapt to evolving attacks. Recent works have demonstrated that reinforcement learning can be used to train stronger attacker and defender models in tandem by applying PPO-style self-play and DPO-style online preference optimization. In this work, we explore the efficacy of GRPO in this setting. Co-training can be challenging because it requires jointly optimizing multiple properties of both the attacker and defender. We therefore shape model outputs using multiple LLM judge-based reward channels and compute advantages with GDPO, which prevents any single channel from dominating. Our method uses a curriculum that progresses from attacker-only single-turn and multi-turn training to co-training, where attacker and defender models are updated in alternation. We show that this method produces highly effective and transferable attacks, and that co-trained defenders reach competitive safety while preserving general utility. Through a controlled ablation, we further identify which components of our training pipeline most affect the resulting balance between safety and utility. Finally, we find that GRPO tends to collapse attacker diversity over training and discuss possible ways to address this limitation.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2606.09701 [cs.CL]
  (or arXiv:2606.09701v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2606.09701

arXiv-issued DOI via DataCite

Submission history

From: Blake Bullwinkel [view email]
[v1] Mon, 8 Jun 2026 16:21:36 UTC (809 KB)
[v2] Wed, 7 Oct 2026 22:55:53 UTC (957 KB)

来源:arXiv:cs.CL · arxiv.org