arXiv:cs.AI· J. de Curt\`o, I. de Zarz\`a·· 6 小时前AI 评分35
LLM 智能体在信息物理系统中的规划策略战略性评估
Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems
AI 导读
研究提出一个受控基准,用四种编码执行器(预定义、顺序、分层、搜索)在辐射状馈线上控制 40 个产消者的需求响应,LLM 仅声明或建议类型化策略。基于 Llama-3.3-70B 的实验显示,强制搜索在五个基线种子中均为最优解;注入目标替换使模式一致率保持 1.0,但累计电压缺口增加 2.68 倍。
正文
Abstract:LLM-agent evaluations commonly measure task success or agreement with a declared plan. In strategic cyber-physical systems, an architecture must also remain appropriate after autonomous participants respond and physics constrains outcomes. We introduce a controlled benchmark of planning-induced control trajectories: ordered planning operations and directives linking execution architecture to strategic response and physical consequences. Four coded executors (predefined, sequential, hierarchical, and search) control demand response for 40 prosumers on a radial feeder. The LLM declares or advises typed policies and mediates communication; schedules, base prosumer dynamics, stochastic actions, and power flow remain explicit code. Paired forced-mode counterfactuals, exact-prompt caching, common response draws with separate randomness streams, critic isolation, and event-level feasibility isolate comparisons. The Llama-3.3-70B experiments on this feeder distinguish three properties. First, forced search is the oracle in all five baseline seeds under the specified objective. Second, injected objective substitution preserves mode agreement at 1.0 while increasing cumulative voltage shortfall by 2.68x. Third, the 144-scenario, 576-episode factorial bank, using three repeated seeds, contains feasible oracles from predefined, sequential, and search. The prespecified stress-held-out ridge has mean regret 90.7 and no observed value over fixed sequential. A post-hoc constraint-aware analysis reduces regret to 29.0; a simple deadline rule attains 28.7, so this gain does not establish a learning advantage. An all-feasible ablation does not improve over fixed search. These are simulation-internal, descriptive comparisons. A five-model, 300-declaration extension tests interface behaviour, not cross-backbone physical rankings; shared-endpoint latency tails motivate probabilistic live feasibility.
| Subjects: | Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI); Systems and Control (eess.SY) |
| Cite as: | arXiv:2608.04265 [cs.MA] |
| (or arXiv:2608.04265v2 [cs.MA] for this version) | |
| https://doi.org/10.48550/arXiv.2608.04265 arXiv-issued DOI via DataCite |
Submission history
From: J. De Curtò [view email]
[v1]
Tue, 4 Aug 2026 22:52:21 UTC (127 KB)
[v2]
Tue, 6 Oct 2026 15:55:01 UTC (127 KB)
来源:arXiv:cs.AI · arxiv.org