跳到正文
arXiv:cs.LG· Subham Rath, Raj Dandekar, Rajat Dandekar, Sreedath Panat·· 4 小时前AI 评分46

pass@k 无法衡量什么:评估后训练后的多样性与能力保留

What pass@k Cannot Measure: Evaluating Diversity and Capability Retention after Post-Training

AI 导读

论文指出 pass@k 仅依赖问题被正确采样的概率,无法反映正确解的分布差异。在 Qwen2.5-1.5B-Instruct 上用 GRPO 和 RFT 做数学后训练,三种多样性指标朝相反方向变化(三个随机种子无重叠),但 pass@8 和 pass@32 在 GSM8K 上无一致赢家,仅低 k 时能区分。

正文

View PDF HTML (experimental)

Abstract:pass@$k$, the fraction of problems a model solves within $k$ sampled attempts, is the field's default protocol for deciding whether reinforcement-learning (RL) post-training on verifiable rewards improved a model. At the population level, pass@$k$ depends only on a problem's probability of a correct sample, with no term for how it is distributed across outputs. We show this gap is not academic. Training Qwen2.5-1.5B-Instruct on grade-school math with Group Relative Policy Optimization (GRPO) and with rejection-sampling fine-tuning (RFT, training on the model's own shortest verifier-passed rollout) moves three complementary diversity measures (token-level entropy, answer-level entropy, unique answers per prompt) in opposite directions, with zero overlap across three seeds per arm. The gap survives restricting to verifier-correct completions only (lexical diversity among correct solutions is 15% lower for GRPO, after controlling for length) and a count-controlled check isolating diversity among incorrect answers alone, ruling out that GRPO's higher accuracy alone explains it. Yet pass@8 and pass@32 show no consistent winner on GSM8K, and a hard MATH-500 subset shows the same pattern: separation only at low $k$. Compared against the starting checkpoint, no trained arm significantly improves hard-problem coverage: RFT is significantly worse, while GRPO is statistically indistinguishable from it - so GRPO's pass@1 edge over RFT reflects a smaller loss relative to Base, not a capability gain, a missing-control issue, not a failure of pass@$k$. On GSM8K, only pass@1, with no role in detecting diversity by construction, separates the arms cleanly, rewarding the arm whose correct solutions are least diverse. We argue this is a concrete instance of a standard evaluation protocol missing a property it is routinely used to certify.
Comments: 10 pages, 2 figures. Accepted to the NeurIPS 2026 Workshop on Transitioning from Pre-training to Post-training (non-archival)
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.07405 [cs.LG]
  (or arXiv:2610.07405v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07405

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Subham Rath [view email]
[v1] Mon, 5 Oct 2026 21:16:32 UTC (53 KB)

来源:arXiv:cs.LG · arxiv.org