arXiv:cs.AI· Xiaohong Chen, David Bucur, Chenglong Ma, Yi Zhang, Lingming Zhang, Sriram Vishwanath, Grigore Rosu·· 4 小时前
评估真实程序中精确输出与检查点状态预测的基准测试
Evaluating Exact Output and Checkpoint-State Prediction in Real Programs
AI 导读
一项新基准测试从源码和输入预测程序最终输出与检查点状态,包含来自 371 个 Python 和 C++ 程序的 400 个案例,在四种模型家族的七个设置下评估,计划预测 11,200 次,其中 11,151 次产生可评分回答。
正文
Abstract:We present a benchmark for predicting final output and checkpoint state from source and input alone. It extends CRUXEval-style output prediction with paired shorter- and longer-trace inputs and checkpoints inside and after a loop. The benchmark contains 400 cases from 371 Python and C++ programs, evaluated under seven settings from four model families without tools or code execution. Of 11,200 planned predictions, 11,151 produced gradable responses. Reasoning-enabled settings outperform their off counterparts by 33.1 to 55.2 percentage points on completed responses. The strongest setting scores 93.0% on shorter-trace final output, 77.0% on longer-trace final output, and 65.5% and 63.5% on the two state tasks; these scores also hold when missing responses count as wrong. Across 2,397 matched Python comparisons with identical source, changing to the longer-trace input yields 528 correct-to-wrong changes and 147 reversals. Source-clustered analyses preserve this accuracy gap, while adjusted Python models give no evidence of a positive incremental association between cumulative state load and error. Changed inputs and checkpoint tasks alter several factors together, so the gaps do not isolate trace length or an internal state-tracking mechanism. The benchmark exposes errors hidden by short-output scores alone.
| Comments: | 15 pages. Accepted at the NeurIPS 2026 Workshop on AI for Verifiable Coding. Includes appendices |
| Subjects: | Software Engineering (cs.SE); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.11889 [cs.SE] |
| (or arXiv:2610.11889v1 [cs.SE] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11889 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xiaohong Chen [view email]
[v1]
Thu, 8 Oct 2026 12:59:07 UTC (437 KB)
来源:arXiv:cs.AI · arxiv.org