跳到正文
arXiv:cs.LG· Ilia Mahrooghi, Aryo Lotfi, Emmanuel Abbe·· 5 小时前AI 评分39

Goldilocks RL:通过调节任务难度逃离稀疏奖励以提升推理能力

Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning

AI 导读

研究者提出 Goldilocks,一种自适应数据选择策略,用 Selector 网络预测每个候选问题在模型多次采样中的奖励标准差,优先选择奖励波动大、即难度适中的问题,并配合 GRPO 训练。在 OpenMathReasoning 和 Polaris 数据集上,Goldilocks 稳定优于标准 GRPO,达到同等性能最多可减少 78% 的优化步数。

正文

View PDF HTML (experimental)

Abstract:Reinforcement learning has emerged as a powerful paradigm for unlocking reasoning capabilities in language models. However, relying on sparse rewards makes this process highly sample-inefficient, as models must navigate vast search spaces with minimal feedback. While classic curriculum learning aims to mitigate this by ordering data based on complexity, prior works have primarily targeted small datasets and do not directly transfer to the large-scale settings typical of modern language model training. Furthermore, the right ordering for a specific model is often unclear. To address this, we propose Goldilocks, an adaptive data-selection strategy that uses a Selector network to predict the standard deviation of rewards across the model's rollouts for each candidate question. The Selector prioritizes questions with high predicted reward variability, corresponding to questions that are neither too easy nor too hard for the model's current capabilities (Goldilocks principle), while training the model with GRPO. By leveraging the model's performance on seen samples, the Selector continuously adapts to the model's evolving abilities. Across the OpenMathReasoning and Polaris datasets, Goldilocks consistently improves over standard GRPO, requiring up to 78% fewer optimization steps to reach the corresponding GRPO performance.
Comments: 42 pages, 23 figures
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2602.14868 [cs.LG]
  (or arXiv:2602.14868v3 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2602.14868

arXiv-issued DOI via DataCite

Submission history

From: Ilia Mahrooghi [view email]
[v1] Mon, 16 Feb 2026 16:01:27 UTC (459 KB)
[v2] Fri, 8 May 2026 13:05:12 UTC (491 KB)
[v3] Fri, 2 Oct 2026 13:05:30 UTC (1,123 KB)

来源:arXiv:cs.LG · arxiv.org