跳到正文
arXiv:cs.AI· Tianyu Chen, Chujia Hu, Dongrui Liu, Xia Hu, Wenjie Wang·· 4 小时前AI 评分60

LPS-Bench:评测计算机使用智能体在长程规划中的安全意识

LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios

AI 导读

研究者发布 LPS-Bench,用于评测计算机使用智能体(CUA)在 MCP 风格工具工作流中长程规划的安全性,覆盖良性请求与对抗诱导两类场景。基准包含 570 个案例,来自 65 个场景、7 个任务域和 9 种规划风险类型,采用模板引导的多智能体流水线生成指令、模拟工具包和案例级安全标准,并经人工审核,无需为每个案例单独搭建应用环境。

正文

View PDF HTML (experimental)

Abstract:Computer-use agents (CUAs) execute multi-stage tasks through tools, where an early unsafe decision can propagate to consequential actions. Evaluating only final outcomes can miss such decisions, while constructing executable environments for new tasks can make benchmark expansion costly. We present LPS-Bench, a benchmark of long-horizon planning safety in MCP-style tool workflows under benign requests and adversarial steering. A template-guided multi-agent pipeline generates user instructions, simulated toolkits, and case-specific safety criteria, followed by human review. This design supports scalable case expansion without building a separate application environment for every test case. LPS-Bench comprises 570 cases derived from 65 scenarios across 7 task domains and 9 planning-risk types, with representative cases additionally adapted to reusable skills. An LLM-based evaluator applies case-specific criteria to complete interaction records, examining tool choices, arguments, and responses to environmental feedback throughout execution. Evaluations of 13 LLM agents reveal persistent failures in both benign and adversarial settings. Prompt-based interventions yield model-dependent gains, but substantial safety failures remain.
Comments: 50 pages. Accepted at NeurIPS 2026, Evaluations and Datasets Track. Revised benchmark description, skill-augmented evaluation, evaluator validation, and appendices
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2602.03255 [cs.AI]
  (or arXiv:2602.03255v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2602.03255

arXiv-issued DOI via DataCite

Submission history

From: Tianyu Chen [view email]
[v1] Tue, 3 Feb 2026 08:40:24 UTC (8,941 KB)
[v2] Fri, 2 Oct 2026 00:39:45 UTC (5,725 KB)

来源:arXiv:cs.AI · arxiv.org