跳到正文
arXiv:cs.AI· Xichen Yan, Chongyang Gao, Kezhen Chen, Guangyi Zhang, Jiaqi Wu, Lixu Wang·· 5 小时前AI 评分35

面向智能体安全误报审计的正例-无标注学习框架

Positive-Unlabeled Learning for Agent Safety False Alarm Auditing

AI 导读

研究者提出一种两阶段正例-无标注(PU)学习框架,用于审计语言模型智能体安全监控的误报,无需误报安全标签且不改动底层监控器。该方法在主流安全监控器上取得 0.6444 的 macro AUPRC,比八个 PU 基线高出 5.27–16.98 个百分点;相比最强基线 PULDA,在 5% 审查预算下多召回 33.3% 的误报。

正文

View PDF HTML (experimental)

Abstract:Safety monitors help safeguard language-model agents interacting with external tools and environments, but conservative monitoring can generate many false alarms, consuming extensive review resources and weakening trust in alerts. Because false and genuine alarms often remain interleaved in native monitor scores, obtaining a reliable cutoff still requires substantial manual verification. In practice, a small set of verified-safe non-alarmed trajectories may be available while alarms remain unlabeled, naturally casting false-alarm auditing as a positive-unlabeled (PU) ranking problem. The key challenge is monitor-induced selection, since observed safe references are accepted by the monitor, while the hidden safe alarms of interest are precisely those it incorrectly flags, making the observed positives poorly representative of the positives to be recovered. To address this challenge, we propose a two-stage framework in which Trust-aware PU Supervision adapts safe references toward the alarm domain and protects plausible false alarms from excessive negative pressure, while Reliability-gated Rank Distillation consolidates consistent ordering preferences from multiple PU reference models into a single student. Consensus-guided Structural Refinement then improves the student ranking using hierarchical safe-reference support, alarm relations, and predicted reference consensus. The framework requires no alarm safety labels for fitting and leaves the underlying monitor unchanged. Across mainstream safety monitors, our method achieves a macro AUPRC of $0.6444$, outperforming eight evaluated PU baselines by 5.27--16.98 absolute percentage points; compared with PULDA, the strongest evaluated PU baseline, it recovers 33.3% more false alarms at a 5% review budget.
Comments: 19 pages
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.02925 [cs.AI]
  (or arXiv:2610.02925v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.02925

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Lixu Wang [view email]
[v1] Fri, 2 Oct 2026 07:16:33 UTC (1,017 KB)

来源:arXiv:cs.AI · arxiv.org