arXiv:cs.LG· Haoxiang Zhang, Qinglin Chen, Hiroaki Hayashi, Zhuofeng Li, Siming Zhang, Jiaxin Zhang, Jixuan Chen, Fang Wu, Pan Lu, Silvio Savarese, Julian McAuley, Chien-Sheng Wu·· 6 小时前AI 评分48
Self-Retrospection Distillation:把事后经验变成事前预判
Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
AI 导读
Self-Retrospection Distillation(SRD)将完整轨迹中的事后经验蒸馏为同一策略在交互前的预判预测,预判仅作训练目标、推理时无需显式生成。
正文
Authors:Haoxiang Zhang, Qinglin Chen, Hiroaki Hayashi, Zhuofeng Li, Siming Zhang, Jiaxin Zhang, Jixuan Chen, Fang Wu, Pan Lu, Silvio Savarese, Julian McAuley, Chien-Sheng Wu
Abstract:Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails. We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting? We introduce prospective learning, which uses post-hoc experience to supervise foresight predictions from the pre-interaction view, and instantiate it with Self-Retrospection Distillation (SRD). Intuitively, a completed trajectory reveals knowledge that would have been useful and pitfalls that should be avoided; SRD distills this privileged hindsight into trajectory-blind foresight of the same policy. Foresight serves only as a training target and need not be explicitly generated at inference time. Across 10 tool-integrated reasoning and long-horizon agentic tasks, SRD complements RLVR and self-distillation baselines with gains of up to $24.2$ pp. Its advantage is especially pronounced when reward contrast is scarce: when $37$--$98\%$ of rollout groups are reward-uniform across model scales, yet SRD can still exploit learning signal from sampled trajectories. In the 2B setting, where $98\%$ of groups are all-failure, the RLVR training ends up at $0.0\%$ success, while adding SRD reaches $60.6\%$ under the same rollout budget. Our results suggest that post-hoc agent experience is useful not only for evaluating or improving behavior, but also for shaping predictive representations before available interaction.
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.08077 [cs.AI] |
| (or arXiv:2610.08077v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08077 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Haoxiang Zhang [view email]
[v1]
Tue, 6 Oct 2026 10:06:47 UTC (6,703 KB)
来源:arXiv:cs.LG · arxiv.org