跳到正文
arXiv:cs.AI· Lindsey Ferris, Sierra Bonilla·· 6 小时前AI 评分45

如何保障 AI 编写的软件:一个医疗平台智能体案例研究

A Case Study in Assuring AI-Written Software

AI 导读

一项案例研究记录了一个由编码智能体构建、由无正式软件工程训练的操作者治理的生产级医疗平台。其工作流演变为人类主导的元智能体系统:一个智能体写代码,其他智能体监督审查,项目规则沉淀经验。

正文

View PDF HTML (experimental)

Abstract:Software-engineering agents can enable people without formal software training to build systems they could not otherwise implement and simultaneously can produce more code than even experts can meaningfully inspect. In both cases, exhaustive code review is not reliable as the sole basis for human control. We report a case study of a production healthcare platform built through coding agents and governed by an operator without formal software-engineering training. Over time, its workflow grew into a human-led meta-agent system where one agent wrote code, other agents supervised and reviewed it, and project rules carried lessons forward. The operator found that tests, monitors and reviewing agents used to supervise the system were fallible. Some monitors measured proxies rather than outcomes, some audits failed silently, missing checks disappeared from reported results and one automated repair caused operational disruption. In this case, human control depended on keeping the intended outcome, the evidence used to judge it, the agents' permissions and the final human decision were all tied to the same underlying objective.
Comments: Accepted to the NeurIPS 2026 Meta-Agents Workshop
Subjects: Software Engineering (cs.SE); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
Cite as: arXiv:2610.08651 [cs.SE]
  (or arXiv:2610.08651v1 [cs.SE] for this version)
  https://doi.org/10.48550/arXiv.2610.08651

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sierra Bonilla [view email]
[v1] Tue, 6 Oct 2026 16:37:08 UTC (190 KB)

来源:arXiv:cs.AI · arxiv.org