arXiv:cs.AI· Jorge Garc\'ia-Carrasco, Sergio Garc\'ia-Carrasco, Alejandro Mat\'e, Juan Trujillo·· 3 小时前
面向嵌入式软件开发的 LLM 智能体闭环评测基准
Closed-loop evaluation of LLM agents for embedded software development
AI 导读
研究者提出一套面向嵌入式编码智能体的闭环评测基准,包含五个嵌入式控制任务和四种反馈场景,实现基于模拟 ESP32 固件,用于检验智能体在传感、时序与安全约束下的闭环行为而非静态代码质量。
正文
Abstract:Large language models (LLMs) are increasingly deployed as coding agents that edit files, run builds and tests, inspect execution results, and repair software iteratively. Embedded firmware is a demanding target because correctness depends on closed-loop behavior under sensing, timing, and safety constraints, not only on static source quality. Yet embedded-agent evaluation remains limited and often emphasizes one-shot synthesis or offline correctness.
We present a benchmark for closed-loop evaluation of embedded coding agents. Each task provides a plain-text engineering description, constrained workspace, and visible build-and-runtime surface. The agent must translate requirements into implementation and self-verification steps, then iterate until the required device behavior is achieved. The suite contains five embedded-control tasks and four feedback scenarios: one-shot generation, realistic self-verification, CI-style red/green feedback, and oracle-style detailed feedback.
The implementation targets simulated ESP32 firmware for reproducibility. We evaluate seven GPT-family and Qwen-family configurations across five tasks and four scenarios, with three repetitions per condition for 420 runs. gpt-5.4 has the highest pass rate among evaluated configurations but does not saturate the benchmark; qwen3.5-27B is the strongest observed local model; and smaller local models degrade sharply in pass rate and search efficiency. These results suggest that capable local embedded coding agents are emerging.
| Comments: | Published in Journal of Systems Architecture 179 (2026) 103937. 23 pages, 4 figures, 5 tables. Code and artifacts: this https URL |
| Subjects: | Software Engineering (cs.SE); Artificial Intelligence (cs.AI); Hardware Architecture (cs.AR); Machine Learning (cs.LG) |
| ACM classes: | D.2.5; I.2.7 |
| Cite as: | arXiv:2610.11447 [cs.SE] |
| (or arXiv:2610.11447v1 [cs.SE] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11447 arXiv-issued DOI via DataCite (pending registration) |
|
| Journal reference: | J. Syst. Archit. 179 (2026) 103937 |
| Related DOI: | https://doi.org/10.1016/j.sysarc.2026.103937
DOI(s) linking to related resources |
Submission history
From: Jorge García-Carrasco [view email]
[v1]
Thu, 8 Oct 2026 08:02:10 UTC (59 KB)
来源:arXiv:cs.AI · arxiv.org