arXiv:cs.LG· Jiaxin Zhang, Xiangyu Peng, Qinglin Chen, Yu Li, Hiroaki Hayashi, Chien-Sheng Wu·· 3 小时前AI 评分42
Prospective Hindsight:用预测-现实差距实现自校准强化学习
Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps
AI 导读
研究者提出 Prospective Hindsight(PH)训练原则,利用智能体行动前的预测与反馈后的评估之间的差距,通过 stop-gradient 惊讶加权优势放大自模型最不准确样本的梯度贡献。
正文
Abstract:Reinforcement learning for long-horizon agents relies on purely retrospective training signals: credit is assigned only after observing environmental consequences, leaving the agent's belief at action time invisible to the gradient. We introduce Prospective Hindsight (PH), a self-calibrating training principle that augments any retrospective base method with a signal derived from the gap between the agent's prospective prediction (before feedback) and the retrospective evaluation (after feedback). This per-rollout surprise identifies samples where the agent's self-model is most inaccurate and amplifies their gradient contribution through a stop-gradient surprise-weighted advantage. Since the prospective predictor shares parameters with the policy, the two co-evolve, progressively shifting focus to the agent's remaining blind spots. We connect this principle to a privileged-information gap and show that minimizing the surprise residual provides a descent pathway on the agent's miscalibration rate; calibration thus emerges as a byproduct of optimization rather than from an added objective. On single-turn verifiable tasks and a multi-turn personal-agent task (under GRPO, on-policy distillation, and their combination), PH improves both task performance and calibration, with consistent gains across model scales. Notably, the dominant miscalibration mode shifts structurally between regimes, overconfident failures in single-turn, underconfident successes in multi-turn, yet the same training principle addresses both successfully.
| Comments: | NeurIPS 2026 |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.02740 [cs.LG] |
| (or arXiv:2610.02740v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02740 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jiaxin Zhang [view email]
[v1]
Fri, 2 Oct 2026 03:12:00 UTC (674 KB)
来源:arXiv:cs.LG · arxiv.org