arXiv:cs.AI· Jisheng Dang, Yushuo Zhao, Dewei Liu, Junfeng Fang, Bimei Wang, Tiantian Rao, Hong Peng, Bin Hu, Tat-Seng Chua·· 5 小时前AI 评分34
DNAlign:面向 LLM 的动态零空间安全对齐框架
DNAlign: Dynamic Null-Space Safe Alignment for LLMs
AI 导读
DNAlign 是一个轻量级 LLM 安全对齐框架,将控制论优化与零空间投影结合,把扰动限制在与有害内容相关的子空间中,从而在降低有害输出的同时保留通用知识与生成质量。该框架用人类偏好数据训练的价值函数自适应优化控制信号,在多个 LLM 基座上的评测中整体表现优于既有对齐基线,且未牺牲生成多样性。代码已开源。
正文
Abstract:Ensuring the safe and reliable deployment of large language models (LLMs) remains a fundamental challenge. Existing safety alignment approaches either incur high computational cost or unintentionally disrupt the model's core knowledge, leading to degraded fluency and factual accuracy on benign tasks. This reveals a persistent trade-off between safety and utility. We propose DNAlign, a lightweight alignment framework that integrates control-theoretic optimization with null-space projection. By treating the LLM as a dynamic system, the proposed framework introduces controllable perturbations to steer generation toward safe behavior. A key component is the projection module, which restricts these perturbations to the harmful-related subspace derived from neutral hidden states, thereby preserving general knowledge and response quality. A value function trained on human preference data adaptively optimizes the control signals to align with human safety preferences. Extensive evaluations across multiple LLM backbones demonstrate that our framework consistently reduces harmful outputs while maintaining fluency, coherence, and factual utility. It achieves superior overall performance compared to prior alignment baselines without sacrificing generation diversity. These results indicate that the proposed framework provides an effective and practically deployable solution for safe LLM alignment. Code is available at this https URL.
| Comments: | Regular Paper; 13 pages, 8 figures, and 1 table |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.02844 [cs.AI] |
| (or arXiv:2610.02844v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02844 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jisheng Dang [view email]
[v1]
Fri, 2 Oct 2026 05:32:49 UTC (4,661 KB)
来源:arXiv:cs.AI · arxiv.org