跳到正文
arXiv:cs.AI· Jorge Garc\'ia-Carrasco, Sergio Garc\'ia-Carrasco, Alejandro Mat\'e, Juan Trujillo·· 3 小时前

面向嵌入式软件开发的 LLM 智能体闭环评测基准

Closed-loop evaluation of LLM agents for embedded software development

AI 导读

研究者提出一套面向嵌入式编码智能体的闭环评测基准,包含五个嵌入式控制任务和四种反馈场景,实现基于模拟 ESP32 固件,用于检验智能体在传感、时序与安全约束下的闭环行为而非静态代码质量。

正文

View PDF HTML (experimental)

Abstract:Large language models (LLMs) are increasingly deployed as coding agents that edit files, run builds and tests, inspect execution results, and repair software iteratively. Embedded firmware is a demanding target because correctness depends on closed-loop behavior under sensing, timing, and safety constraints, not only on static source quality. Yet embedded-agent evaluation remains limited and often emphasizes one-shot synthesis or offline correctness.
We present a benchmark for closed-loop evaluation of embedded coding agents. Each task provides a plain-text engineering description, constrained workspace, and visible build-and-runtime surface. The agent must translate requirements into implementation and self-verification steps, then iterate until the required device behavior is achieved. The suite contains five embedded-control tasks and four feedback scenarios: one-shot generation, realistic self-verification, CI-style red/green feedback, and oracle-style detailed feedback.
The implementation targets simulated ESP32 firmware for reproducibility. We evaluate seven GPT-family and Qwen-family configurations across five tasks and four scenarios, with three repetitions per condition for 420 runs. gpt-5.4 has the highest pass rate among evaluated configurations but does not saturate the benchmark; qwen3.5-27B is the strongest observed local model; and smaller local models degrade sharply in pass rate and search efficiency. These results suggest that capable local embedded coding agents are emerging.
Comments: Published in Journal of Systems Architecture 179 (2026) 103937. 23 pages, 4 figures, 5 tables. Code and artifacts: this https URL
Subjects: Software Engineering (cs.SE); Artificial Intelligence (cs.AI); Hardware Architecture (cs.AR); Machine Learning (cs.LG)
ACM classes: D.2.5; I.2.7
Cite as: arXiv:2610.11447 [cs.SE]
  (or arXiv:2610.11447v1 [cs.SE] for this version)
  https://doi.org/10.48550/arXiv.2610.11447

arXiv-issued DOI via DataCite (pending registration)

Journal reference: J. Syst. Archit. 179 (2026) 103937
Related DOI: https://doi.org/10.1016/j.sysarc.2026.103937

DOI(s) linking to related resources

Submission history

From: Jorge García-Carrasco [view email]
[v1] Thu, 8 Oct 2026 08:02:10 UTC (59 KB)

来源:arXiv:cs.AI · arxiv.org