跳到正文
arXiv:cs.AI· Zheng Fang, Dongming Jin, Yihong dong, Yongmin Li, Kechi Zhang, Zhi Jin, Ge Li·· 5 小时前AI 评分43

ClarifyCodeBench:评估 LLM 澄清代码生成模糊需求的能力

ClarifyCodeBench: Evaluating LLMs on Clarifying Ambiguous Requirements for Code Generation

AI 导读

研究者推出 ClarifyCodeBench,一个基于真实编程任务构建的交互式基准,用于评估 LLM 澄清需求歧义的能力,并配套人工标注的歧义类型、澄清问题与标准答案。

正文

View PDF HTML (experimental)

Abstract:Large Language Models have emerged as programming assistants. However, the efficacy of code generation is constrained by the quality of input requirements, which are frequently ambiguous, incomplete, or underspecified. While LLMs excel at one-shot code synthesis, their ability to proactively clarify intent remains underexplored, as a critical trait for robust software engineering. Existing benchmarks largely overlook this interactive bottleneck, assuming perfectly specified prompts that do not reflect the iterative nature of requirement elicitation. To bridge this gap, we introduce ClarifyCodeBench, a novel interactive benchmark for evaluating LLMs' capability in resolving requirement ambiguity. Constructed from real-world programming tasks, ClarifyCodeBench features high-quality manual annotations, including N unique ambiguity types, associated clarification questions, and corresponding ground-truth answers. Furthermore, we formalize two rigorous metrics to assess the interaction quality: Turn-discounted Key Question Rate, which penalizes inefficient questioning, and Optimal Round Adherence, which measures the precision of the elicitation process. We conduct a systematic evaluation of six state-of-the-art LLMs using ClarifyCodeBench. Our empirical results yield three critical insights: 1) Capability Decoupling: Strong code generation performance does not inherently translate to effective requirement clarification; 2) The Reasoning Paradox: While increased computational thinking enhances code correctness, it yields marginal gains in identifying ambiguities; 3) The Multi-ambiguity Ceiling: LLMs' clarification performance degrades sharply as the density of ambiguities increases, revealing a significant bottleneck in handling complex, real-world specifications. Our work underscores the necessity for future AI4SE research to transition from static synthesis to interactive elicitation.
Comments: Code: this https URL
Subjects: Software Engineering (cs.SE); Artificial Intelligence (cs.AI)
Cite as: arXiv:2607.00711 [cs.SE]
  (or arXiv:2607.00711v2 [cs.SE] for this version)
  https://doi.org/10.48550/arXiv.2607.00711

arXiv-issued DOI via DataCite

Submission history

From: Zheng Fang [view email]
[v1] Wed, 1 Jul 2026 09:58:59 UTC (278 KB)
[v2] Mon, 31 Aug 2026 03:57:35 UTC (278 KB)

来源:arXiv:cs.AI · arxiv.org