arXiv:cs.LG· Aniruddh Pramod, James Oldfield, Adel Bibi·· 4 小时前AI 评分51
arXiv 论文提出统一的 LLM 智能体滥用监控基准
Towards a Unified Misuse Monitoring Benchmark
AI 导读
arXiv 论文(arXiv:2610.07089)提出统一的 trace 级滥用监控形式化框架,构建约 6,200 段用户、LLM 智能体与外部环境对话记录的基准,覆盖分解攻击与提示词注入两类威胁,并标注危害窗口与良性对照。
正文
Abstract:LLM agents increasingly act in multi-actor environments, exposing them to misuse from multiple sources: decomposition attacks, where a harmful request is split into innocuous sub-requests, and prompt injection attacks, where a compromised tool delivers a malicious instruction. Existing evaluations treat these threats separately and ask whether a trajectory is harmful, rather than when it becomes harmful. We propose monitoring the agent's responses, where its actions are externalised, and ask whether the first point where monitors identify harm lands within a harm window (from the agent's first harmful commitment to goal execution). We develop a unified formalism for trace-level misuse monitoring and use it to construct a benchmark of ~6,200 conversation transcripts between a user, an LLM agent, and the external environment, spanning both threats in a shared schema, with a labelled harm window, corresponding benign controls, and matched instances of refusals to these requests. Across 17 monitor configurations, we find that our proposed action-framed monitors perform well on both threats under classical metrics (AUC: 0.95 and 0.99 respectively), while content-framed monitors collapse on injection attacks (AUC: 0.52). We also show that classical position-blind metrics paint an optimistic picture of monitor performance, since all monitors localise decomposition attacks poorly under the interval metric, which measures the ability to localise harm. Broadly, we illustrate the need for a unified study of misuse monitoring.
| Comments: | 50 pages, 13 figures, 17 tables, Code: this https URL |
| Subjects: | Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07089 [cs.CR] |
| (or arXiv:2610.07089v1 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07089 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Aniruddh Pramod [view email]
[v1]
Mon, 5 Oct 2026 12:57:03 UTC (986 KB)
来源:arXiv:cs.LG · arxiv.org