跳到正文
arXiv:cs.AI· Yunji Kim, Yunseok Lee, Hyunwoo Seo, Jaerim Choi, Woojin Lee·· 5 小时前AI 评分48

PLCWorld:在闭环工厂仿真中评测 LLM 生成的 PLC 程序

PLCWorld: Benchmarking LLM-Generated PLC Programs in Closed-Loop Plant Simulation

AI 导读

研究者推出 PLCWorld,一个将结构化文本(ST)执行与仿真工厂响应、传感器反馈耦合的闭环执行环境与基准,含 100 个合成任务和 473 个任务-条件对,覆盖运动控制与物料搬运。直接使用 GPT-5.5 在 Easy 用例上任务成功率为 82.70%,Hard 用例仅 25.10%。代码、仿真环境、基准任务与基线实现均已公开。

正文

View PDF HTML (experimental)

Abstract:Programmable logic controllers (PLCs) coordinate industrial equipment by reading sensor inputs and issuing control commands. Evaluating whether large language model (LLM)-generated PLC programs satisfy task requirements and safety constraints requires observing how their commands affect device and workpiece states. We introduce PLCWorld, a common closed-loop execution environment and benchmark that couples Structured Text (ST) execution with simulated plant responses and sensor feedback. Grounded in control relations identified in industrial PLC programs and engineering documentation, PLCWorld contains 100 synthetic tasks and 473 registered task-condition pairs across Motion Control and Material Handling, with difficulty defined by control-dependency scope. A common protocol reports Task Success and Safety Violation separately. Validation combines practitioner review, reference and alternative programs, targeted counterexamples, specification-evaluator alignment checks, and comparisons with independent ST runtimes. Reference and alternative programs satisfy their applicable cases, while all 542 targeted counterexamples activate their designated evaluator rules under at least one registered condition. Execution Gap relates submission-profile acceptance to subsequent task failure or observed Safety Violation. Across the constructed task groups, direct GPT-5.5 achieves 82.70% Task Success on Easy cases but 25.10% on Hard cases. Evaluations of six LLMs and four adapted generation-and-verification workflows further expose differences between completion, safety, and generation cost. Our code, simulation environment, benchmark tasks, and baseline implementations are publicly available at this https URL.
Comments: 36 pages, 9 figures. Project website: this https URL
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.02982 [cs.AI]
  (or arXiv:2610.02982v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.02982

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Woojin Lee [view email]
[v1] Fri, 2 Oct 2026 08:14:42 UTC (7,644 KB)

来源:arXiv:cs.AI · arxiv.org