arXiv:cs.AI· Hanyu Wang, Nakul Agarwal, Hossein Nourkhiz Mahjoub, Ehsan Moradi Pari, Makoto Fukushima, Jinghui Chen, Vaishnav Tadiparthi·· 4 小时前
推理强化学习中轨迹 rollout 的参考引导与自由生成如何平衡:Adaptive Reference Guidance 方法
Balancing Reference Guidance and Free Generation in Trajectory Rollouts for Reasoning RL
AI 导读
研究提出 Adaptive Reference Guidance(ARG),通过前缀续写从续写结果中学习一个跨题目共享的前缀选择器,无需估计成功概率或额外生成,在固定生成预算内构造正确轨迹,并应用于 GRPO 的全失败组。在 Qwen3-4B 和 Qwen3-8B 上,ARG 在五个数学推理基准中取得最高的 aggregate pass@12,平均采样准确率也具竞争力。
正文
Abstract:A verified reference solution provides a correct trajectory for training a reasoning model. Alternatively, a prefix of the reference can guide the model in generating a trajectory of its own. How much reference guidance should we provide? We study this question through prefix continuation, where the model continues from a reference prefix and keeps the resulting trajectory if it passes verification, falling back to the reference otherwise. Since both procedures produce correct trajectories, we compare their distributions with the ideal distribution, the model's own distribution conditioned on successful verification. For one continuation, we derive the KL divergence in closed form, which, up to a bounded term, decreases with the product of the probability of generating a different correct trajectory and the reference surprisal, the negative log probability of the reference suffix given the prefix. Since a longer prefix tends to raise the former but lowers the latter, continuation success alone does not determine the preferred amount of guidance. From this analysis, we learn a prefix selector shared across training questions from continuation outcomes, without estimating success probabilities or additional generation. The resulting Adaptive Reference Guidance (ARG) constructs correct trajectories within a fixed generation budget, and we apply it to all-failure groups in Group Relative Policy Optimization (GRPO). Experiments on Qwen3-4B and Qwen3-8B across five mathematical reasoning benchmarks show that ARG achieves the highest aggregate pass@12 among the evaluated methods with competitive average sampled accuracy.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.11128 [cs.AI] |
| (or arXiv:2610.11128v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11128 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hanyu Wang [view email]
[v1]
Thu, 8 Oct 2026 02:54:48 UTC (411 KB)
来源:arXiv:cs.AI · arxiv.org