arXiv:cs.AI· Youwei Feng, Yitong Zhang, Yuetong Liu, Jia Li·· 4 小时前
Agent 安全不止防违禁行为:ObligationGuard 与 ObligationBench 识别未履行义务
Safe Actions Alone Do Not Ensure Safe Agents: Identifying Unfulfilled Obligations with Guard Models
AI 导读
研究指出,仅识别被禁止的操作不足以保障 LLM 智能体安全,还需识别未履行的安全关键义务;在 GLM-5.3 轨迹中 56.92% 含未履行义务,仅 30.00% 含违禁操作。
正文
Abstract:Guard models are increasingly used to safeguard LLM-based agents, primarily by identifying actions that agents are forbidden to perform. However, identifying forbidden actions alone is insufficient to ensure agent safety. In this paper, we argue that agent safety also depends on identifying required yet unperformed safety-critical actions, which we call obligations. Our preliminary study on a popular benchmark for evaluating safety shows that 56.92% of GLM-5.3 trajectories contain unfulfilled obligations, compared with only 30.00% containing forbidden actions. This finding reveals unfulfilled obligations as a major and previously overlooked source of safety risk. However, to our knowledge, no existing benchmark evaluates whether guard models can identify these obligations. To close this gap, we introduce ObligationBench, the first benchmark for evaluating the capability of obligation identification, comprising 240 expert-validated trajectories covering issue resolution, feature development, and terminal operations. Our evaluation of 14 representative models reveals substantial limitations: the highest recall and exact-match rate are only 48.97% and 10.00%, respectively. To address these limitations, we develop ObligationGuard using 40,000 synthetic training examples. ObligationGuard achieves 57.52% recall and an exact-match rate of 21.67%, surpassing all evaluated models on both metrics. We call on the community to incorporate obligation identification into the design and evaluation of future guard models to improve agent safety.
| Comments: | 21 pages. Code, data, and supplementary materials: this https URL |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.11773 [cs.AI] |
| (or arXiv:2610.11773v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11773 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Youwei Feng [view email]
[v1]
Thu, 8 Oct 2026 11:54:04 UTC (1,454 KB)
来源:arXiv:cs.AI · arxiv.org