跳到正文
arXiv:cs.CL· Hui Wu, Xiaoyang Wang, Zhong Fan·· 4 小时前AI 评分37

PowerCodeBench:面向 LLM 电力系统代码生成的知识边界探测与需求引导干预

Knowledge boundary probing and demand-guided intervention for LLM-based power system code generation

AI 导读

研究者发布 PowerCodeBench——一个面向 pandapower 的冻结 2,000 任务参数化基准,并提出无需更新权重的部署时工作流,通过文档驱动的 L0-L3 探测生成各模型 API 画像,并用查询侧需求估计器在生成前选择分层 API 证据。

正文

View PDF HTML (experimental)

Abstract:Large language models (LLMs) can turn grid-analysis requests into executable programs for power-system simulation, but utilities and research laboratories often require on-premise deployment. In this setting, first-pass failures frequently arise at an API-knowledge boundary, through hallucinated functions, misused parameters, and mishandled result tables. We present PowerCodeBench, a parameterised benchmark generator released as a frozen 2,000-task suite for pandapower, and a deployment-time workflow that requires no weight updates. Documentation-driven L0-L3 probes produce per-model API profiles for diagnosis, model comparison, documentation allocation, and backend calibration. A query-side demand estimator selects layered API evidence before generation, while execution feedback routes targeted repair. Across ten open-weight LLMs (1.5B-480B) and four mid-tier APIs, the validation-enabled workflow raises scalar-match accuracy by 32-56 percentage points after up to three repair rounds relative to an unassisted first pass, for every model of at least 7B and every API. Open-weight models in the 70B-120B range reach the four-vendor mid-tier accuracy range under matched no-tool conditions. Selective injection approaches the full-layer reference using 41% of its prompt tokens. Among model-item pairs passing numerical checks under both workflows, engineering review confirms the requested analysis in 88% of full-workflow outputs versus 66% under plain repair. Round-0 pilots on OpenDSS and PyPSA motivate staged onboarding from broad retrieval at cold start to calibrated selective injection. Measured throughput, latency, energy, and allocated GPU memory establish a practical on-premise serving envelope.
Comments: 52 pages, 10 figures, including supplementary material. Revised following peer review; expanded validation, cross-backend pilots, serving measurements, and supplementary material. Published in Advanced Engineering Informatics
Subjects: Software Engineering (cs.SE); Computation and Language (cs.CL); Systems and Control (eess.SY)
Cite as: arXiv:2605.31478 [cs.SE]
  (or arXiv:2605.31478v2 [cs.SE] for this version)
  https://doi.org/10.48550/arXiv.2605.31478

arXiv-issued DOI via DataCite

Journal reference: Advanced Engineering Informatics 77 (2027) 105328
Related DOI: https://doi.org/10.1016/j.aei.2026.105328

DOI(s) linking to related resources

Submission history

From: Hui Wu [view email]
[v1] Fri, 29 May 2026 16:06:34 UTC (991 KB)
[v2] Wed, 7 Oct 2026 02:35:18 UTC (916 KB)

来源:arXiv:cs.CL · arxiv.org