arXiv:cs.LG· Chuanyu Qin, Chenxu Yang, Qingyi Si, Naibin Gu, Dingyu Yao, Zheng Lin, Peng Fu, Nan Duan, Jiaqi Wang·· 2 天前AI 评分38
让模型向"近未来"的自己学习:RLVR 的时序自蒸馏方法
Learning from the Near Future: Temporal Self-Distillation for RLVR
AI 导读
研究者提出时序自蒸馏,让策略从自身训练后期更强的 checkpoint 获取指导,并认为最有用的"时序教师"不必是最强的,而需在新增能力与学习者兼容性之间平衡。其中 NPO 用已验证的未来自我轨迹做离策略行为迁移,在八个图文基准上把 GRPO 从 60.25 提升到 62.84,AutoNPO 达 63.15;NPD 的近未来教师在持续 RL 后达 63.23,高于远未来教师的 61.92。
正文
Abstract:Reinforcement learning with verifiable rewards (RLVR) is a core post-training recipe for reasoning models, yet pure on-policy learning can be inefficient when useful trajectories are difficult to discover or exploration narrows. Existing self-guided approaches largely reuse capability already available to the current or earlier learner. We instead ask whether learning can also make use of capabilities that emerge later in training: can a model learn from its own future self? We introduce temporal self-distillation, in which a policy receives guidance from a stronger later checkpoint of itself. We hypothesize that the most useful temporal teacher need not be the strongest one: a teacher must provide sufficiently new capability while remaining compatible enough for that capability to be readily transferred, motivating a near-future regime. We study this principle through two complementary mechanisms. Near-Future Policy Optimization (NPO) performs off-policy behavioral transfer using verified future-self trajectories, while Near-Future Policy Distillation (NPD) performs on-policy token-level transfer on learner-generated trajectories. We further introduce AutoNPO, which adaptively determines when temporal guidance is useful and how far to roll back, turning future-self guidance into a repeated self-bootstrap process. Across eight image-text benchmarks, NPO improves GRPO from 60.25 to 62.84 and AutoNPO reaches 63.15, with consistent gains on text-only and video reasoning. Under NPD, a near-future teacher reaches 63.23 after continued RL versus 61.92 with a far-future teacher, despite lower immediate post-distillation performance. Together, these results suggest that effective temporal self-distillation depends not simply on teacher strength, but on a balance between newly acquired capability and learner compatibility.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2604.20733 [cs.LG] |
| (or arXiv:2604.20733v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2604.20733 arXiv-issued DOI via DataCite |
Submission history
From: Chenxu Yang [view email]
[v1]
Wed, 22 Apr 2026 16:20:41 UTC (10,559 KB)
[v2]
Thu, 1 Oct 2026 01:38:55 UTC (12,985 KB)
来源:arXiv:cs.LG · arxiv.org