arXiv:cs.AI· Qitao Tan, Xiaoying Song, Arman Akbari, Arash Akbari, Yanzhi Wang, Xiaoming Zhai, Lingzi Hong, Zhen Xiang, Jin Lu, Geng Yuan·· 7 小时前AI 评分41
Palette:面向 LLM 的模块化可控按需授权安全对齐放松框架
Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs
AI 导读
针对基础模型"一刀切"拒绝策略导致授权专业人员合法请求被误拒的问题,研究者提出 Palette 框架,通过多目标搜索定位拒绝方向并经轻量适配内化到模型中,可在授权目标域选择性放松拒绝行为。Palette 支持模块化组合,各领域安全控制独立学习后经参数合并拼装,无需重训练即可实现多域按需授权。在四个安全基准、多个模型变体及 LLM 与 VLM 上的实验显示,该框架实现精确安全控制且不牺牲通用效用。
正文
Abstract:Current safety alignment of foundation models largely follows a \emph{one-size-fits-all} paradigm, applying the same refusal policy across users and contexts. As a result, models may refuse requests that are unsafe for general users but legitimate for authorized professionals, limiting helpfulness in specialized professional settings. Existing approaches either require costly realignment or rely on inference-time steering that suffers from imprecise control and added latency. To this end, we propose \textsc{Palette}, a modular, controllable, and efficient framework that selectively relaxes refusal behavior on authorized target domains while preserving standard safety elsewhere. Our method identifies a refusal direction via multi-objective search and internalizes it into the model through lightweight adaptation. \textsc{Palette} further supports modular composition: it learns domain-specific safety controls independently and composes them through parameter merging, enabling on-demand multi-domain authorization without retraining. Experiments across four safety benchmarks, multiple model variants, and both LLMs and VLMs show that \textsc{Palette} delivers precise safety control without sacrificing general utility, offering a practical path toward foundation models that adapt to diverse professional needs.
| Subjects: | Artificial Intelligence (cs.AI); Software Engineering (cs.SE) |
| Cite as: | arXiv:2605.24154 [cs.AI] |
| (or arXiv:2605.24154v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2605.24154 arXiv-issued DOI via DataCite |
Submission history
From: Qitao Tan [view email]
[v1]
Fri, 22 May 2026 19:22:17 UTC (16,759 KB)
[v2]
Mon, 5 Oct 2026 17:57:43 UTC (16,760 KB)
来源:arXiv:cs.AI · arxiv.org