arXiv:cs.AI· Jingnan Zheng, Dongcheng Zhang, Yi Zhang, Ming Zhang, Qiaosheng Zhang, Youbang Sun, An Zhang, Xiangnan He, Tat-Seng Chua, Xia Hu, Bowen Zhou, Chaochao Lu, Xiang Wang·· 3 小时前
ReSI:面向抗攻击与韧性 AI 的递归安全改进框架
ReSI: Recursive Safety Improvement toward Resistant and Resilient AI
AI 导读
研究团队提出递归安全改进框架 ReSI,通过自动化研究多轮迭代红队测试与模型更新,提升模型对已知威胁的抗性与对未知风险的韧性。在四个稠密与 MoE 模型上,ReSI 将 X-Teaming 平均攻击成功率从 86.01% 降至 31.45%,优于 GPT-5.6-Luna 的 56.69%。
正文
Authors:Jingnan Zheng, Dongcheng Zhang, Yi Zhang, Ming Zhang, Qiaosheng Zhang, Youbang Sun, An Zhang, Xiangnan He, Tat-Seng Chua, Xia Hu, Bowen Zhou, Chaochao Lu, Xiang Wang
Abstract:Recursive self-improvement, the participation of AI systems in improving their own capabilities, is beginning to move from theoretical prospect to practice, posing both challenges and opportunities for safety alignment. Models evolve through frequent updates, and their safety alignment requires continual adaptation to each new checkpoint. Meanwhile, with evolving red-teaming methods exposing new vulnerabilities, safety improvement for each checkpoint needs to mitigate exposed vulnerabilities and generalize to risks not yet revealed. Following R$^2$AI, we term these goals resistance to known threats and resilience to unforeseen risks. Recursive self-improvement, in turn, inspires an approach to both goals: safety alignment could likewise advance through successive rounds of evaluation and update. We therefore introduce ReSI, a recursive safety improvement framework that implements this approach through automated research. In each round, ReSI applies diverse red-teaming methods to identify vulnerabilities in the current target model, develops training recipes, and promotes the update with the largest safety gain among those passing a Pareto gate on capability retention as the next target model. Across four dense and mixture-of-experts models, ReSI matches or exceeds evaluated frontier models on in-distribution and out-of-distribution safety benchmarks, and outperforms alignment baselines on nearly all safety evaluations while largely preserving general capabilities. In particular, ReSI reduces the mean X-Teaming attack success rate across the four models from 86.01% to 31.45%, well below GPT-5.6-Luna's leading frontier result of 56.69%, indicating stronger resilience to attacks unseen during training. These findings support recursive safety improvement as a practical path toward resistant and resilient AI.
| Subjects: | Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.12233 [cs.CR] |
| (or arXiv:2610.12233v1 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12233 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jingnan Zheng [view email]
[v1]
Thu, 8 Oct 2026 16:17:21 UTC (1,668 KB)
来源:arXiv:cs.AI · arxiv.org