arXiv:cs.AI· Xingru Zhou, Luis Sentis, Aarti Choudhary·· 6 小时前AI 评分40
SAFESHIELD:面向小语言模型部署时安全的决策组织框架
SAFESHIELD: A Decision-Organization Framework for Deployment-Time Safety of Small Language Models
AI 导读
SAFESHIELD 是面向小语言模型部署时安全的决策组织框架,将安全决策拆分为准入、路由、证据、放行四类职责,并把已提交决策记录为可审计的 Decision Traces,该工作已被 IEEE TPS 2026 应用赛道接收。
正文
Abstract:Deployment-time safety of language models is commonly implemented through runtime guardrails such as input moderation, routing, retrieval verification, and output filtering. Existing deployment frameworks provide increasingly capable mechanisms for these functions, but offer limited guidance on how the safety decisions they produce should be explicitly organized, coordinated, and audited. We formulate deployment-time safety as a decision-organization problem with two elements: responsibility-oriented decomposition of safety decisions and explicit coordination among them. We instantiate this formulation in SAFESHIELD, a deployment-time safety system for small language models that organizes four recurring decision responsibilities (admission, routing, evidence, and release) and records committed decisions in auditable Decision Traces. We evaluate SAFESHIELD through mechanism-level experiments, aggregate stage ablations, controlled coordination ablations, and a deployment-oriented stress suite. Mechanism-level results show that the instantiated safeguards provide the capabilities required by the decision process, while aggregate ablations show substantial degradation in end-to-end safety as the surrounding safety organization is removed. More importantly, dedicated coordination ablations preserve the participating safeguard mechanisms while selectively severing their dependencies: removing admission gating substantially increases false release, and withholding upstream evidence from the release decision reduces release accuracy from 96.0% to 69.5%. These results provide system-level evidence that deployment-time safety depends not only on the capability of individual guardrails, but also on how their decisions are organized and coordinated.
| Comments: | Accepted to the Application Track of IEEE TPS 2026. 12 pages |
| Subjects: | Software Engineering (cs.SE); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR) |
| Cite as: | arXiv:2610.07276 [cs.SE] |
| (or arXiv:2610.07276v1 [cs.SE] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07276 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xingru Zhou [view email]
[v1]
Mon, 5 Oct 2026 19:16:05 UTC (1,233 KB)
来源:arXiv:cs.AI · arxiv.org