arXiv:cs.LG· Xiaohan Yi, Wen Luo, Yani Huang, Junfeng Zhan, Asher Qin, Peilin Zhao, Xi Xiao·· 9 小时前AI 评分45
GC-OPD:基于图条件的在策略智能体蒸馏,用现成教师模型提升多轮任务成功率
Graph-Conditioned On-Policy Agent Distillation from Off-the-Shelf Teachers
AI 导读
研究者提出 Graph-Conditioned On-Policy Agent Distillation(GC-OPD),通过图索引重复的教师执行轨迹,为现成教师模型的评分补充执行证据,无需任务特定的教师优化。
正文
Abstract:On-policy distillation (OPD) trains compact language agents with teacher feedback on student-generated trajectories. In multi-turn tasks, compounding errors can move students beyond the teacher's effective supervision. We introduce Graph-Conditioned On-Policy Agent Distillation (GC-OPD), which enriches an off-the-shelf teacher's scoring context with execution evidence. A graph indexes repeated teacher executions by shared states while preserving complete successful and failed histories. After each student episode, GC-OPD retrieves current-state references or historical alternatives and combines them with student hindsight to score the original thought-action tokens. Using the same original teachers, GC-OPD improves mean success over vanilla OPD from 24.70% to 48.78% on ScienceWorld (4B student), from 53.36% to 85.26% on ALFWorld Unseen, and from 29.10% to 37.65% on WebShop. At matched student sizes, it also achieves higher mean success than every evaluated OPD baseline using GRPO-trained teachers on ScienceWorld and ALFWorld; the strongest such ScienceWorld 4B baseline reaches 46.66%. GC-OPD requires no task-specific teacher optimization.
| Comments: | 18 pages, 3 figures |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.37522 [cs.LG] |
| (or arXiv:2609.37522v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.37522 arXiv-issued DOI via DataCite |
Submission history
From: Xiaohan Yi [view email]
[v1]
Tue, 29 Sep 2026 13:12:44 UTC (325 KB)
[v2]
Fri, 2 Oct 2026 06:22:53 UTC (325 KB)
来源:arXiv:cs.LG · arxiv.org