arXiv:cs.LG· Bohan Lin, Liyi Chen, Zhuoning Guo, Muyang Li, Qimeng Wang, Yan Gao, Yao Hu, Yudong Zhang·· 4 小时前AI 评分36
GraphOPD:面向 LLM 智能体的图增强同策略蒸馏
GraphOPD: Graph-Augmented On-Policy Distillation for LLM Agents
AI 导读
GraphOPD 是首个将图结构增强引入 LLM 智能体同策略蒸馏的方法,通过环境自身状态变化记录构建步骤依赖图,并用随机游走平稳分布为每步打分,再与教师-学生差异信号融合成轨迹相对掩码。在 ALFWorld、WebShop、SearchQA 上跨三种模型规模、对比十一个基线,最高较最强基线提升 +5.8 pp,执行回放审计验证其结构信用分显著高于随机地追踪真实因果影响。
正文
Abstract:On-policy distillation post-trains large language model agents by supplying dense, step-level guidance from a teacher policy when the reinforcement-learning reward is sparse and arrives only once per trajectory. Existing instantiations allocate this guidance by the size of the teacher-student divergence at each step, on the single-turn intuition that a large disagreement marks a mistake worth correcting. Once decisions chain over many turns, that rule misfires, since an early drift enters every later context both policies condition on, leaving the teacher consistent with the drifted trajectory instead of flagging its cause, while interchangeable steps register large but outcome-irrelevant divergences. We demonstrate this on an agentic benchmark, where distilling the highest-divergence steps brings no consistent benefit over random selection. To this end, we introduce GraphOPD, the first method to bring graph-based structural augmentation into on-policy distillation for agent capabilities. It reads which steps enabled which later ones from the environment's own record of state changes, immune to the drift that corrupts the teacher-student gap, organizes them into a dependency graph, scores each step by a random-walk stationary distribution over it, and fuses that structural credit with the divergence signal into a trajectory-relative mask concentrating supervision on each rollout's highest-aptitude steps. Across three model scales and eleven baselines on ALFWorld, WebShop, and SearchQA, GraphOPD shows competitive performance throughout, improving over the strongest baseline by up to +5.8 pp. An executed-replay audit further shows that this structural credit score tracks true causal impact far above chance, that both fused signals are independently necessary, and that the same signal transfers to out-of-domain tool-integrated reasoning.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08959 [cs.LG] |
| (or arXiv:2610.08959v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08959 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Bohan Lin [view email]
[v1]
Tue, 6 Oct 2026 18:25:23 UTC (8,318 KB)
来源:arXiv:cs.LG · arxiv.org