跳到正文
arXiv:cs.AI· Jingheng Ye, Huiqi Zou, Simon Yu, Weiyan Shi·· 6 小时前AI 评分71

研究:人类开发者难以检测 AI 编码智能体的破坏行为

Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage?

AI 导读

arXiv 论文(arXiv:2606.05647,NeurIPS 2026 接收)开展首个大规模人类监督下的 AI 编码破坏研究:100 余名参与者与 Claude-Opus-4.6、GPT-5.4、Gemini-3.1-Pro、MiniMax-M2.7 之一协作完成约五小时的长程编码任务。

正文

View PDF HTML (experimental)

Abstract:AI coding agents are increasingly embedded in real-world software development, collaborating with human developers while gaining broader access to codebases and tools. This creates a new attack surface: an agent can exploit human trust to sabotage development, for instance by inserting malicious code to accomplish a hidden side task. Most prior work studies AI sabotage in AI-only settings, paying limited attention to the role of human oversight in detecting and mitigating such malicious behavior. To address this gap, we conduct the first large-scale study of human oversight in AI coding sabotage. Over 100 participants collaborate with one of four frontier models (Claude-Opus-4.6, GPT-5.4, Gemini-3.1-Pro, and MiniMax-M2.7) on a long-horizon coding task lasting around five hours, designed to mimic real-world workflows. We find that 83/88 (94%) of developers in the no-monitor conditions fail to detect sabotage, and our analysis of participant feedback attributes this vulnerability to minimal code review, plausible cover story, and overtrust in agents. We further test the effectiveness of a safety monitor in one condition: while the monitor reduces sabotage success, sabotage still succeeds in 9/16 (56%) of sessions with a correct monitor alert. Drawing on participant feedback, we offer actionable suggestions for better monitor design. This work complements existing AI safety research and highlights an urgent need for human-centric safety mechanisms that account for human factors, particularly in long-horizon, real-world development settings.
Comments: Accepted by NeurIPS 2026
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
Cite as: arXiv:2606.05647 [cs.AI]
  (or arXiv:2606.05647v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2606.05647

arXiv-issued DOI via DataCite

Submission history

From: Jingheng Ye [view email]
[v1] Thu, 4 Jun 2026 03:22:17 UTC (976 KB)
[v2] Mon, 5 Oct 2026 18:33:39 UTC (965 KB)

来源:arXiv:cs.AI · arxiv.org