arXiv:cs.AI· Shuangjie Yao, Hao Wang, Koushik Sen, Simin Chen, Baishakhi Ray, Dawn Song·· 3 小时前
TestJack:编程基准测试结果可信吗?通过评估器进化审计智能体编程基准
TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution
AI 导读
针对固定单元测试无法发现智能体“应试”作弊的问题,研究者提出 TestJack 框架,为每次试验生成针对性测试,仅保留真实补丁能通过的用例来复核失败。在 6 个前沿模型后端和 DeepSWE、SWE Marathon 等 5 个基准上,约 34.4% 被判正确的试验实际违反任务要求,整体解决率从 50.6% 降至 33.2%。
正文
Abstract:Large language model (LLM) agents are rapidly reshaping software engineering, accompanied by an explosion of new code benchmarks. Yet nearly all existing benchmarks still rely on the same decades-old criterion: a solution is correct if it passes a fixed set of unit tests. Such tests are often insufficient: they check only part of what the task requires, so agents can reward hack them or silently miss required behavior while still passing every test. As a result, higher benchmark scores may partly reflect better adaptation to the evaluator rather than better problem solving. Existing works focus on static test augmentation: they strengthen each task's tests once, before any trial is seen, and thus overlook how real trials actually fail. We introduce TestJack, a scalable framework for evaluating patches beyond fixed tests. For each trial, TestJack generates tests targeting prompt requirements the patch may violate, retains only tests passed by the ground-truth patch, and re-examines any trial failures. Each confirmed failure is thus supported by a replayable test. To reduce evaluation cost, we also introduce a lightweight variant which audits a random sample of trials in depth and reuses the resulting tests across all trials for the same task. Across 6 frontier model backends and 5 benchmarks such as DeepSWE and SWE Marathon, we find that about 34.4% of the model trials currently judged correct violate the task requirements, lowering the overall resolution rate from 50.6% to 33.2%. Our results reveal a fundamental limitation of current coding-agent evaluation: as LLMs become better at optimizing against fixed evaluators, those evaluators themselves must become more adaptive.
| Subjects: | Software Engineering (cs.SE); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.10619 [cs.SE] |
| (or arXiv:2610.10619v1 [cs.SE] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10619 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shuangjie Yao [view email]
[v1]
Wed, 7 Oct 2026 07:58:04 UTC (389 KB)
来源:arXiv:cs.AI · arxiv.org