arXiv:cs.AI· Ravishka Rathnasuriya, Wei Yang·· 4 小时前
LLM 代码生成的过度自信失败问题研究
Characterizing Overconfident Failure in LLM-Based Code Generation
AI 导读
研究揭示 LLM 代码生成存在"过度自信失败":错误程序在 token 级置信度上与正确程序难以区分,覆盖四个开源代码模型和三个基于执行的 benchmark。现有不确定性信号仅提供部分且依赖模型的执行失败证据,基于不确定性的筛选无法稳定提升接受集准确率。指令微调会增加失败生成的确定性却不改善正确性判别,常见缓解手段也无法可靠解决该问题。
正文
Abstract:Large language models (LLMs) are increasingly used for automated code generation, but generated programs can appear syntactically plausible while still failing execution-based correctness checks. Existing validation methods, such as testing and program analysis, remain essential but are often incomplete, costly, or applied only after generation. Model-derived uncertainty is therefore a natural early reliability signal. This paper studies the dilemma of overconfidence in code LLMs where incorrect programs are often generated with token-level confidence comparable to correct programs. We study this dilemma across four open-source code models and three execution-based benchmarks. Our analysis begins by investigating whether existing uncertainty metrics provide reliable proxies for execution correctness in code generation. We then characterize overconfidence at both global and local token levels, asking whether incorrect programs remain indistinguishable from correct ones under confidence and entropy summaries, including selective generation and the limits of instruction tuning. Finally, we evaluate whether common mitigation strategies reduce this failure mode. Our study yields four findings. First, existing uncertainty signals provide only partial and model-dependent evidence of execution failure. Second, overconfidence persists at both program and token levels, and uncertainty-based selection does not consistently improve accepted-set accuracy. Third, instruction tuning can increase certainty on failing generations without consistently improving correctness discrimination. Fourth, common mitigation techniques improve specific aspects of reliability but do not reliably resolve overconfident failure. Our exploratory latent analysis suggests that hidden representations may encode correctness-related signals that output confidence does not expose.
| Subjects: | Software Engineering (cs.SE); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.11300 [cs.SE] |
| (or arXiv:2610.11300v1 [cs.SE] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11300 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ravishka Rathnasuriya [view email]
[v1]
Thu, 8 Oct 2026 06:07:44 UTC (879 KB)
来源:arXiv:cs.AI · arxiv.org