跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Feiyu Gavin Zhu, Qi Xu, Zhifei Deng, Zhigang Hua, Luke Simon, Jean Oh, Reid Simmons·· 14 小时前AI 评分33

通过语义 rollout 分析迭代优化策略结构

Iterative Policy Refinement through Semantic Rollout Analysis

AI 导读

研究者提出一种闭环框架,利用 LLM 对策略 rollout 的语义分析迭代优化结构化策略,无需人工指令即可识别并修正策略结构中的次优之处。在赛车和开门任务上,该方法比零样本 LLM 生成的结构将模仿学习性能提升最多 15%,达到同等强化学习性能所需算力减少 75%。结果表明,表格化 rollout 分析可作为有效反馈信号,使 LLM 生成的策略结构与专家演示对齐。

正文

View PDF HTML (experimental)

Abstract:Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the expert demonstrations. We propose a closed-loop framework that iteratively refines structured policies using LLM-guided analysis of policy rollouts. By logging rollouts as semantically meaningful tabular data and prompting the LLM to generate diagnostic analysis code, our method identifies suboptimalities in the policy structure and iteratively corrects them without requiring human instruction. Experiments on car racing and door opening tasks show that our approach improves imitation learning performance by up to 15% over zero-shot LLM-generated structures and requires 75% less compute to achieve the same reinforcement learning performance. These results demonstrate that tabular rollout analysis provides an effective feedback signal to align LLM-generated policy structures with expert demonstrations, and we can utilize it to generate good policy structures automatically.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.01652 [cs.LG]
  (or arXiv:2610.01652v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01652

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Feiyu Zhu [view email]
[v1] Thu, 1 Oct 2026 13:15:42 UTC (1,636 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org