arXiv:cs.AI· Aleksi Huotala, Miikka Kuutila, Mika M\"antyl\"a·· 3 小时前
LLM 在软件工程系统综述筛选中的性能是否已停滞?
Has LLM Screening Performance Stalled in Software Engineering Systematic Reviews?
AI 导读
研究基于 SESR-Eval 基准评估八款新 LLM 的论文筛选性能,平均 MCC 仅从 0.347 升至 0.365,提升有限。不同系统综述之间的差异仍大于不同 LLM 之间的差异,LLM 尚未能取代人工筛选,新模型的高成本带来的优势不明显。
正文
Abstract:Screening in systematic reviews (SRs) is manual and time-consuming. Prior work has explored large language models (LLMs) for automating this step, but LLMs are evolving rapidly, so earlier performance claims may no longer accurately reflect their screening performance. We used an existing software engineering SR screening benchmark (SESR-Eval) as our data. We also power-sampled a new, smaller dataset (SESR-Eval-Mini) that allows evaluation at lower costs. Using this data, we evaluated eight new LLMs for screening performance. Additionally, we tested different prompts, analyzed LLM agreement in screening decisions and criteria, and examined the effect of refining the inclusion and exclusion criteria on screening performance. The eight new LLMs performed marginally better than the seven old ones: avg. MCC across secondary studies rose from 0.347 to 0.365. Differences between secondary studies are still bigger than between LLMs. Computing the overall screening decision from criterion-level decisions degraded screening performance only slightly. LLMs generally agree with each other in their corresponding screening decisions (mean Gwet's AC1 = 0.830), though certain inclusion and exclusion criteria showed larger disagreement than others. Refining the inclusion and exclusion criteria slightly improved recall and made decisions easier for some LLMs, but overall impacts of criteria refinement were modest. LLMs are not yet ready to replace humans in paper screening and the advantages new, more costly models bring, appear to be very limited. Agent-based approaches, prompt engineering, and further criteria refinement are three potential future research avenues.
| Comments: | 53 pages, four external figures available in the research artifact |
| Subjects: | Software Engineering (cs.SE); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.10633 [cs.SE] |
| (or arXiv:2610.10633v1 [cs.SE] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10633 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Aleksi Huotala [view email]
[v1]
Wed, 7 Oct 2026 13:25:54 UTC (196 KB)
来源:arXiv:cs.AI · arxiv.org