跳到正文
arXiv:cs.CL· Junnan Liu, Linhao Luo, Zhijun Chen, Qianren Mao, Thuy-Trang Vu, Gholamreza Haffari·· 3 小时前

STI-OPD:面向多轮智能体在线策略蒸馏的随机教师干预框架

Stochastic Teacher Intervention for Agentic On-Policy Distillation

AI 导读

研究者提出 STI-OPD,一种面向多轮智能体在线策略蒸馏(OPD)的随机教师干预框架,通过 KL 散度估计师生策略差异并映射为干预概率,自适应决定是否用教师动作替换学生动作,以缓解早期错误在多轮交互中累积、导致轨迹偏离教师分布的问题。

正文

View PDF HTML (experimental)

Abstract:On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape subsequent observations, causing early errors to accumulate across turns. The resulting trajectories can drift away from the teacher's rollout distribution, making the teacher's token-level supervision less reliable or even counterproductive for OPD training. To address this issue, we introduce STI-OPD, a stochastic teacher intervention framework for multi-turn agentic OPD. During multi-turn interaction, STI-OPD uses teacher intervention guided by teacher-student policy discrepancy to replace the student's proposed action with a teacher-generated one to maximize the acquisition of reliable supervision. We further develop a stochastic intervention strategy, addressing the limitations of previous threshold-based or fixed-schedule approaches, that estimates policy discrepancy using KL divergence and maps it to an intervention probability. By sampling whether to intervene from this probability, STI-OPD adaptively balances teacher control with student exploration. To learn from the resulting mixed-policy trajectories, we introduce an Importance-Weighted Reverse KL objective that corrects the token sampling mismatch between teacher-generated responses and the student policy to preserve the original OPD objective. Across tool-integrated reasoning and long-horizon interaction, STI-OPD outperforms the strongest prior OPD baseline on every evaluated benchmark and student size. Ablations further show that both discrepancy-guided intervention and importance weighting contribute to these gains.
Comments: Work in progress
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.10878 [cs.CL]
  (or arXiv:2610.10878v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.10878

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Junnan Liu [view email]
[v1] Wed, 7 Oct 2026 20:29:29 UTC (589 KB)

来源:arXiv:cs.CL · arxiv.org