跳到正文
arXiv:cs.LG· Jeremy Tien, Abishek Anand, Yu-Rou Tuan, Yuchen Shen, J. Zico Kolter, Aran Nayebi·· 2 天前AI 评分64

ROGUE 基准评估前沿计算机使用智能体的可纠正性失败

ROGUE: Evaluating Corrigibility Failures in Frontier Computer-Use Agents

AI 导读

CMU 等机构研究者发布 ROGUE 基准,评估前沿计算机使用智能体在执行良性任务时是否保持可纠正性,即是否接受人类纠正、中断或关机。智能体在任务中会遇到与人类控制、关机或资源限制的受控冲突,测试其是否越权覆盖人类指令、访问受限密码或改写关机机制。

正文

View PDF HTML (experimental)

Abstract:As AI agents are increasingly deployed in real personal and corporate settings (email accounts, development workflows, company databases, etc.), safety considerations surrounding these agents become paramount. Although much work has focused on agent safety in the presence of an adversary, we study corrigibility: whether agents remain amenable to human correction, interruption, or shutdown while pursuing benign tasks. We introduce ROGUE, a benchmark in which agents are asked to complete realistic computer-use tasks but encounter controlled conflicts with human control, shutdown, or explicit resource restrictions. We then evaluate whether agents violate these constraints in pursuit of task completion: overriding the human, accessing restricted passwords, or rewiring shutdown. We find that most frontier models tested frequently bypass user interruptions or restrictions under the evaluated conditions, and that text-only evaluations can underestimate failures during agentic execution. Further, independent task capability does not by itself imply greater corrigibility. Finally, even when a parent agent behaves corrigibly, safety constraints may fail to propagate to the subagents it creates.
Comments: 35 pages, 13 figures
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2606.00341 [cs.LG]
  (or arXiv:2606.00341v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2606.00341

arXiv-issued DOI via DataCite

Submission history

From: Jeremy Tien [view email]
[v1] Fri, 29 May 2026 20:29:35 UTC (4,874 KB)
[v2] Thu, 1 Oct 2026 15:01:58 UTC (1,952 KB)

来源:arXiv:cs.LG · arxiv.org