arXiv:cs.AI· Utku Boran Torun, Veli Karakaya, Eray T\"uz\"un·· 4 小时前
LLM 能否无需代码修复 Bug?面向 no-code 修复的自动化验证研究
Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes
AI 导读
一项研究提出基于执行的自动化流水线,评估 LLM 在真实浏览器环境中生成 no-code 修复的能力,测试了此前基准发布的 12 种配置生成的 322 个修复。
正文
Abstract:A no-code fix resolves an invalid bug report by directing the user to change a setting, update to a version where the problem is already fixed, or adjust their workflow. Manually verifying whether a proposed no-code fix resolves the reported bug takes considerable developer time. This study proposes an automated, execution-based pipeline for evaluating the capability of large language models (LLMs) to generate no-code fixes in a real browser environment. We evaluate 322 no-code fixes generated by the 12 configurations released with the benchmark of a previous study, covering bug reports categorized as Faulty Configuration, Wrong Version, or External System & Dependency. An executor agent applies each fix by following its natural-language instructions, and an issue-specific checker determines whether the reported bug persists. We repeat the pipeline with three executors: two Computer-Use Agents, OpenCUA-72B and Claude Sonnet 5, and one multimodal agentic LLM, Meta's Muse Glimmer. Only 17.6% of the candidate issues could be set up and passed both sanity gates. Across the 322 fixes, 14.6% to 49.7% resolved the bug depending on the executor, and the strongest configuration, Claude Opus 4.6 in the Vanilla pipeline, resolved up to 74.1% of its fixes under Claude Sonnet 5. Changing only the executor shifted a configuration's resolution rate by 38.8% on average, and the three executors reached the same verdict on only 46.9% of the fixes. Compared with human execution, the executors matched the human consensus for 66.1% to 88.1% of the sampled fixes. Even under the best executor, fewer than half of the LLM-generated no-code fixes resolve the reported bug, so such fixes need verification before they reach users. Execution-based verification can provide this, but the measured capability depends strongly on the executor, which evaluations must report and control.
| Subjects: | Software Engineering (cs.SE); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.11963 [cs.SE] |
| (or arXiv:2610.11963v1 [cs.SE] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11963 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Utku Boran Torun [view email]
[v1]
Thu, 8 Oct 2026 13:44:53 UTC (344 KB)
来源:arXiv:cs.AI · arxiv.org