跳到正文
arXiv:cs.LG· Anvi Kalpesh Shah, Umamaheswara Sharma B·· 6 小时前AI 评分35

超越排行榜:稠密与 MoE 模型在自动程序修复中的多维度评估

Beyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program Repair

AI 导读

研究提出基于 ISO/IEC 25010 的加权质量指数(QI),综合功能正确性、可维护性、安全性和生成效率评估自动程序修复模型。

正文

View PDF HTML (experimental)

Abstract:Automated Program Repair (APR) with language models is usually evaluated by whether a generated patch passes the test suite, which can hide differences in maintainability, security, and computational cost. We propose a Weighted Quality Index (QI), inspired by the ISO/IEC 25010 software quality model, that combines functional correctness, maintainability, security, and generation efficiency under configurable weighting schemes. We evaluate three dense Qwen2.5-Coder models (3B, 7B, 14B) and the 16B-parameter DeepSeek-Coder-V2-Lite Mixture-of-Experts (MoE) model (2.4B active parameters) on 40 QuixBugs and 90 Defects4J bugs, all run locally on identical hardware to control for infrastructure effects. Model rankings change with the weighting scheme, showing that single-metric evaluation can hide trade-offs. The MoE model shows almost no statistically significant difference in correctness from the 7B and 14B dense models (McNemar's exact test) while using 3-6 times fewer active parameters, whereas correctness increases significantly across the three dense scales. These results suggest that active parameter count can be a more informative lens than total parameter count for sparse code models.
Subjects: Software Engineering (cs.SE); Machine Learning (cs.LG)
Cite as: arXiv:2610.08173 [cs.SE]
  (or arXiv:2610.08173v1 [cs.SE] for this version)
  https://doi.org/10.48550/arXiv.2610.08173

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Anvi Kalpesh Shah [view email]
[v1] Tue, 6 Oct 2026 11:27:12 UTC (244 KB)

来源:arXiv:cs.LG · arxiv.org