跳到正文
arXiv:cs.CL· Fagun Patel, Sang T. Truong, Duc Q. Nguyen, Kazunori Fukuhara, Benjamin W. Domingue, Sanmi Koyejo, Nick Haber·· 4 小时前AI 评分38

CodeInsight:面向迭代式问题求解建模的大规模数据集

A Dataset for Modeling Iterative Problem-Solving

AI 导读

研究者构建了 CodeInsight 数据集,包含 2 门 C++ 入门课程、3286 名本科生在 2 个学年内的超 300 万次代码提交,带测试用例级结果、时间戳与源码。基于该数据集搭建的基准在统一校准与评分协议下对比参数化、序列与生成式模型,适配的 RSSM 在 4 门课程中的 3 门取得最高预测准确率,LLM 预测器准确率较低但能生成完整提交。模型编码能力与该场景下的预测表现呈反比。

正文

View PDF HTML (experimental)

Abstract:Solving problems through repeated attempts is a sequential modeling task: at each step, the solver receives feedback and decides how to revise their solutions. Predicting whether performance improves, plateaus, or regresses across attempts is central to understanding any iterative problem-solving process in both human learners and autonomous agents. Beyond outcomes, modeling what errors persist and how strategies shift across attempts provides deeper insight into the mechanics of sequential learning. Studying these dynamics requires observing many solvers as they attempt, receive feedback, and revise. Programming courses with automated grading provide this setting, as students iteratively submit code to test suites and receive feedback on every attempt. We therefore curate CodeInsight, a large-scale dataset of over 3 million submissions from 3,286 undergraduates across 2 introductory C++ courses in 2 academic years, with test-case-level outcomes, timestamps, and source code. On this dataset, we build a benchmark that evaluates models spanning parametric, sequential, and generative traditions under a shared calibration-and-scoring protocol, including a Recurrent State Space Model (RSSM) adapted to track solver characteristics through discrete latent variables and an LLM-based predictor that generates explicit solutions. The adapted RSSM achieves the strongest predictive accuracy on three of the four courses. The LLM predictor is less accurate but produces full submissions at each attempt, enabling direct analysis of failure modes. We find that the model's coding proficiency is inversely related to predictive performance in this setting, with the LLM better understood as a generative solver conditioned on context rather than a faithful predictor of solver behavior. We publicly release our code and the dataset on request to facilitate future research.
Comments: EMNLP 2026 Findings
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2609.00940 [cs.CL]
  (or arXiv:2609.00940v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.00940

arXiv-issued DOI via DataCite

Submission history

From: Quang Duc Nguyen [view email]
[v1] Tue, 1 Sep 2026 09:00:55 UTC (8,387 KB)
[v2] Wed, 7 Oct 2026 17:58:22 UTC (8,387 KB)

来源:arXiv:cs.CL · arxiv.org