arXiv:cs.AI· Tianyu Chen, Chujia Hu, Wenjie Wang·· 4 小时前
Safety Sentry:通过 EXECUTE-ASK-REFUSE 路由实现上下文感知的人工干预
SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing
AI 导读
Safety Sentry 将 LLM agent 动作安全防护重构为 {EXECUTE, ASK, REFUSE} 三路逐实例路由决策,其推理仅需一次解码调用。单一解码时阈值可让同一 checkpoint 在不同风险容忍度部署中重新定位而无需重训练,该轻量 guard model 在整体准确率与安全召回上超越多种开放权重与前沿闭源基线,同时控制双向错误率。
正文
Abstract:LLM agents act on real-world environments through tool calls, and a single misjudged action can cause irreversible harm. The standard safeguard is a guard model that labels each proposed action as safe or unsafe, but this binary view conflates two distinct decisions: whether the action is harmful in itself, and whether it is appropriate given the user's context. It also operates at the granularity of action categories rather than individual instances, producing routine interruptions that erode autonomy and train users to wave through the most consequential alerts. We reframe the problem as a per-instance three-way routing decision over {EXECUTE, ASK, REFUSE} and instantiate it with Safety Sentry, a lightweight guard model whose inference reduces to a single decoding call. A single decoding-time threshold lets one fixed checkpoint be re-positioned across deployments of differing risk tolerance without retraining. Safety Sentry outperforms a broad set of open-weight and frontier closed-source baselines on overall accuracy and safety-related recall, while controlling both directional error rates this http URL and data are available at this https URL
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.13594 [cs.AI] |
| (or arXiv:2607.13594v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.13594 arXiv-issued DOI via DataCite |
Submission history
From: Chujia Hu [view email]
[v1]
Wed, 15 Jul 2026 08:38:16 UTC (1,220 KB)
[v2]
Thu, 8 Oct 2026 06:32:32 UTC (1,042 KB)
来源:arXiv:cs.AI · arxiv.org