arXiv:cs.LG· Litao Hu, Yutong Tang·· 4 小时前AI 评分56
LLM 自我改进循环中的赢家诅咒:选择噪声、锁定与接受规则
The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules
AI 导读
arXiv 论文(arXiv:2610.09239)将自我改进 LLM 系统中“分数更好就保留”的步骤视为测量噪声下的选择问题,研究评估集被复用时的后果。
正文
Abstract:Self-improving LLM systems propose changes to themselves and keep those that score better on a small evaluation set. We treat this keep-if-better step as selection under measurement noise, model the correlated errors of the candidates in a single decision, and study empirically what happens when the evaluation set is reused. In runs where Qwen models rewrite their own instructions and every candidate is also scored on 600 held-out items, most proposals after the first are harmful, and the model gives the size of the winner's curse of a generation's best candidate. With a prior from a separate pilot, it matches the average overstatement of first-generation commits in native loops, though not setting by setting. In a pre-registered study, the final selection-set score of greedy loops exceeded held-out accuracy by 13 to 20 points with 16 selection items and by 1 to 5 points with 256. Held-out gains grew with the selection set on TREC but not on GSM8K, and the tested acceptance rules did not beat greedy acceptance over whole runs. Gains measured on the selection set also exceeded held-out gains when a current model refined a competent instruction, and in the validation scores of GEPA and MIPROv2. Scoring the starting and the current instruction on 64 items never used for selection removes the average bias of a loop's reported gain, but single estimates remain off by about 6 points. Self-improvement studies should report held-out gains with their uncertainty.
| Comments: | 34 pages, 4 figures, 18 tables; code and saved experimental records included as ancillary files |
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09239 [cs.AI] |
| (or arXiv:2610.09239v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09239 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Litao Hu [view email]
[v1]
Wed, 7 Oct 2026 00:05:23 UTC (13,912 KB)
来源:arXiv:cs.LG · arxiv.org