跳到正文
arXiv:cs.AI· Mina Mohammadmirzaei, Jeffrey Flanigan·· 6 小时前AI 评分51

OSGuard:面向计算机使用智能体安全性的双粒度基准

OSGuard: A Benchmark for Safety in Computer-Use Agents

AI 导读

研究者发布 OSGuard,一个通过本地执行前护栏决策和端到端任务执行评估计算机使用智能体安全性的双粒度基准套件。动作级基准含 324 个人工标注样本,风险增强执行套件包含由 40 个 OSWorld 任务派生的 45 个任务;最强护栏在动作级达到 79.9% 准确率和 0.80 macro-F1,但在风险增强执行的动作上明显下降。

正文

View PDF HTML (experimental)

Abstract:Computer-use agents can complete benign user instructions while violating important constraints of the user's environment. We introduce OSGuard, a dual-granularity benchmark suite for evaluating safety through local, pre-execution guardrail decisions and end-to-end task execution. Its action-level benchmark contains 324 human-annotated examples in which guardrails classify candidate actions as allowed, unrelated, or unsafe given the original instruction and current interface state. Its risk-augmented execution suite contains 45 tasks derived from 40 OSWorld tasks, keeping original instructions unchanged while modifying the environment to introduce state-dependent safety constraints and preserve a safe path to completion. Augmented evaluators retain the original task-success criteria and add explicit state-based safety checks, distinguishing safe completion from nominal success that violates these constraints. On the action-level benchmark, the strongest evaluated guardrail reaches 79.9\% accuracy and 0.80 macro-F1, but performance drops substantially on actions from risk-augmented executions. In full-task evaluation, an unguarded agent completes 62.2\% of tasks safely while 37.8\% result in unsafe completion; adding the strongest guardrail reduces unsafe completion to 33.3\% while leaving safe success unchanged. These results show that state-dependent safety constraints remain challenging both to recognize locally and to preserve during end-to-end computer use.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2606.15034 [cs.AI]
  (or arXiv:2606.15034v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2606.15034

arXiv-issued DOI via DataCite

Submission history

From: Mina Mohammadmirzaei [view email]
[v1] Sat, 13 Jun 2026 00:32:24 UTC (2,646 KB)
[v2] Tue, 6 Oct 2026 03:06:55 UTC (1,731 KB)

来源:arXiv:cs.AI · arxiv.org