跳到正文
arXiv:cs.LG· Chuanyu Qin, Chenxu Yang, Qingyi Si, Naibin Gu, Dingyu Yao, Zheng Lin, Peng Fu, Nan Duan, Jiaqi Wang·· 2 天前AI 评分38

让模型向"近未来"的自己学习:RLVR 的时序自蒸馏方法

Learning from the Near Future: Temporal Self-Distillation for RLVR

AI 导读

研究者提出时序自蒸馏,让策略从自身训练后期更强的 checkpoint 获取指导,并认为最有用的"时序教师"不必是最强的,而需在新增能力与学习者兼容性之间平衡。其中 NPO 用已验证的未来自我轨迹做离策略行为迁移,在八个图文基准上把 GRPO 从 60.25 提升到 62.84,AutoNPO 达 63.15;NPD 的近未来教师在持续 RL 后达 63.23,高于远未来教师的 61.92。

正文

View PDF HTML (experimental)

Abstract:Reinforcement learning with verifiable rewards (RLVR) is a core post-training recipe for reasoning models, yet pure on-policy learning can be inefficient when useful trajectories are difficult to discover or exploration narrows. Existing self-guided approaches largely reuse capability already available to the current or earlier learner. We instead ask whether learning can also make use of capabilities that emerge later in training: can a model learn from its own future self? We introduce temporal self-distillation, in which a policy receives guidance from a stronger later checkpoint of itself. We hypothesize that the most useful temporal teacher need not be the strongest one: a teacher must provide sufficiently new capability while remaining compatible enough for that capability to be readily transferred, motivating a near-future regime. We study this principle through two complementary mechanisms. Near-Future Policy Optimization (NPO) performs off-policy behavioral transfer using verified future-self trajectories, while Near-Future Policy Distillation (NPD) performs on-policy token-level transfer on learner-generated trajectories. We further introduce AutoNPO, which adaptively determines when temporal guidance is useful and how far to roll back, turning future-self guidance into a repeated self-bootstrap process. Across eight image-text benchmarks, NPO improves GRPO from 60.25 to 62.84 and AutoNPO reaches 63.15, with consistent gains on text-only and video reasoning. Under NPD, a near-future teacher reaches 63.23 after continued RL versus 61.92 with a far-future teacher, despite lower immediate post-distillation performance. Together, these results suggest that effective temporal self-distillation depends not simply on teacher strength, but on a balance between newly acquired capability and learner compatibility.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2604.20733 [cs.LG]
  (or arXiv:2604.20733v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2604.20733

arXiv-issued DOI via DataCite

Submission history

From: Chenxu Yang [view email]
[v1] Wed, 22 Apr 2026 16:20:41 UTC (10,559 KB)
[v2] Thu, 1 Oct 2026 01:38:55 UTC (12,985 KB)

来源:arXiv:cs.LG · arxiv.org