arXiv:cs.AI· Jingyu Zhang, Shruti Palaskar, Daniel Khashabi, Benjamin Van Durme, Leon A. Gatys, Joseph Yitan Cheng·· 6 小时前AI 评分57
SIGMA:基于 Model Spec 的安全对齐自我改进方法
SIGMA: Self-Improving Alignment Generalization from a Model Spec
AI 导读
arXiv 论文 arXiv:2610.07935 提出 SIGMA,一个仅凭 Model Spec 即可实现安全对齐自我改进的数据生成与训练流程,通过 spec 引导的任务合成和以模型自身为奖励模型的 SFT 与 rubric 强化学习完成训练。
正文
Abstract:LLM agents are increasingly capable of executing complex tasks and of recursively improving themselves on easy-to-verify objectives such as software engineering and mathematics. Since alignment is much harder to verify, this creates a growing risk of capabilities increasing without appropriate safety alignment, especially as capabilities expand to auto-research and cybersecurity. Existing approaches focus on capability self-improvement using verifiable feedback or on alignment training with supervision from stronger models or curated data, creating an external supervision bottleneck for alignment. We ask whether current models can improve their own safety alignment, and propose SIGMA, a data generation and training pipeline enabling alignment self-improvement that generalizes to out-of-distribution settings. Given only a "Model Spec" stating the model's desired behavior, SIGMA leverages a model's reasoning capabilities to strengthen its own safety reasoning. SIGMA first performs spec-guided task synthesis, using the candidate model as a task designer agent to generate diverse alignment dilemma scenarios and convert them into training tasks that stress-test its understanding of the Model Spec. Next, SIGMA conducts self-judged alignment training through supervised fine-tuning and rubric-based reinforcement learning with the model itself as the reward model. Despite training only on single-turn chat data, SIGMA improves safety alignment in multi-turn agentic environments (AgentHarm harmfulness decreases from 22.6 to 14.8; Agentic Misalignment decreases from 79.1 to 3.8), outperforms Deliberative Alignment and Constitutional AI baselines, and retains general capability. Analyses show that a Model Spec balancing harmlessness and helpfulness, test-time reasoning for safety deliberation, and high-quality rubrics from SIGMA's task designer agent are crucial for effective self-improvement.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07935 [cs.AI] |
| (or arXiv:2610.07935v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07935 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jingyu Zhang [view email]
[v1]
Tue, 6 Oct 2026 08:09:56 UTC (14,846 KB)
来源:arXiv:cs.AI · arxiv.org