arXiv:cs.AI· Lars Simon, Holger Eble, Manuel Radons·· 3 小时前
通过回看来学习规划:用于训练推理模型的 Hindsight 层级方法
Learning to Plan by Looking Back: Hindsight Hierarchies for Training Reasoning Models
AI 导读
该研究提出一种推理模型自我改进循环:即使问题难度超出模型当前求解能力,额外提供的解答仍可让模型事后提取有用的解题思路。模型被联合训练以具备三项能力:仅从问题预测解题思路、从问题与已知解答逆向工程思路、以及利用给定思路解题,循环在逆向工程与联合训练间交替。作者给出了方法的形式化规范,并在 Lean 定理证明器中实例化用于交互式定理证明,实证评估留待未来工作。
正文
Abstract:We introduce a self-improvement loop for reasoning models based on the following observation: Even when the difficulty of a problem exceeds the model's current solving abilities, an additionally supplied solution might enable the model to extract useful solution ideas in hindsight. We operationalize this by jointly training the same model to exhibit the following three capabilities: predicting solution ideas from problems alone, reverse-engineering ideas from problems and known solutions, and solving problems using provided ideas. The loop alternates between reverse engineering such ideas from problems with supplied solutions and using these ideas as additional supervision for joint training of all three capabilities. We give a formal specification of our method and a concrete instantiation for interactive theorem proving in the Lean theorem prover; empirical evaluation remains future work.
| Comments: | 21 Pages, 4 Figures |
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Logic in Computer Science (cs.LO) |
| MSC classes: | 68V15, 68T05, 68V20, 68T20 |
| Cite as: | arXiv:2610.12168 [cs.AI] |
| (or arXiv:2610.12168v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12168 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Manuel Radons [view email]
[v1]
Thu, 8 Oct 2026 15:42:27 UTC (35 KB)
来源:arXiv:cs.AI · arxiv.org