跳到正文
arXiv:cs.LG· Masahiro Kato·· 3 小时前AI 评分41

GTDD:面向 AI 编程智能体的生成式测试驱动开发与对抗测试

GTDD: Generative Test-Driven Development for AI Coding Agents with Adversarial Testing

AI 导读

研究者提出生成式测试驱动开发(GTDD),让独立的测试智能体在每次候选实现固定后,依据人类指定的行为契约与历史反馈生成新输入,由可信评估器校验并回传精简反例用于回归测试。在有状态键值存储的配对实验中,开发过程中重新生成测试的策略平均失败率低于仅由同一语言模型一次性生成测试的策略,且向测试方提供候选源码未带来可检测的额外提升。

正文

View PDF HTML (experimental)

Abstract:Test-driven development gives AI coding agents executable requirements for implementing software. Because these agents can adapt their implementations to the examples they observe, passing a predetermined collection of tests can leave substantial parts of the intended behavior unimplemented. We propose Generative Test-Driven Development (GTDD), a formulation of test-driven development in which a separate testing agent generates new inputs after each candidate implementation is fixed, using a human-specified behavioral contract and the feedback from earlier rounds. A trusted evaluator checks these inputs, returns reduced counterexamples to the coding agent, and saves them for regression testing, so development continually confronts failures beyond the initial examples. We characterize the evidence that this process provides through a finite-population analysis of false acceptance under adaptive candidate selection. The resulting bounds quantify how test visibility and repeated feedback affect acceptance, and show that fresh random audits after candidate commitment control false acceptance across development rounds. In a paired experiment on a stateful key-value store, both policies that regenerated tests during development ended with lower mean failure rates than the policy whose tests were generated once by the same language model, and giving the tester the candidate's source produced no detectable additional improvement. Further conditions requesting equal numbers of tests did not isolate any single feature of the policies as the source of this difference. GTDD combines this adaptive development feedback with established regression tests and an independent acceptance rule.
Subjects: Software Engineering (cs.SE); Machine Learning (cs.LG); Methodology (stat.ME); Machine Learning (stat.ML)
Cite as: arXiv:2610.02952 [cs.SE]
  (or arXiv:2610.02952v1 [cs.SE] for this version)
  https://doi.org/10.48550/arXiv.2610.02952

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Masahiro Kato [view email]
[v1] Fri, 2 Oct 2026 07:45:33 UTC (48 KB)

来源:arXiv:cs.LG · arxiv.org