arXiv:cs.AI· NaHyeon Park, Minhyun Lee, Hyunjung Shim·· 5 小时前AI 评分33
CALM:用局部反事实校正替代全局不安全信号,提升文生图安全防护的选择性
Keep It CALM: Analyzing the Limits of Global Unsafety in Text-to-Image Generation
AI 导读
针对文生图免训练安全防护中"全局不安全信号"的假设,研究者提出 CALM(Counterfactual Adaptive Local Modulation),用提示词局部的反事实校正替代统一的全局移除。
正文
Abstract:Training-free safeguards for text-to-image generation often rely on a reusable safety signal, such as an unsafe direction or global toxic subspace, applied broadly across prompts. We provide a controlled geometric analysis of this global-unsafety assumption and reveal a consistent coverage-selectivity trade-off: compact unsafe subspaces fail to cover heterogeneous unsafe semantics, whereas broader aggregation increasingly distorts safety-adjacent benign prompts. Motivated by this finding, we propose CALM (Counterfactual Adaptive Local Modulation), a training-free safeguard that replaces uniform global removal with prompt-local counterfactual correction. Using matched unsafe-benign anchors, CALM routes each prompt to active unsafe categories, minimally edits only violating token representations toward the safe side, and suppresses positively aligned unsafe residual components. Across broad evaluation, CALM significantly improves unsafe content suppression while preserving benign utility, demonstrating that local counterfactual correction provides a more selective alternative to global unsafe signal removal.
| Comments: | NeurIPS 2026 |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.02300 [cs.AI] |
| (or arXiv:2610.02300v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02300 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: NaHyeon Park [view email]
[v1]
Thu, 1 Oct 2026 17:55:02 UTC (8,010 KB)
来源:arXiv:cs.AI · arxiv.org