跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Chiwun Yang, Xiaoyu Li·· 5 小时前AI 评分40

预训练与中训练能从奖励中学到什么?基于 Qwen2.5 的研究

What Pretraining and Midtraining Make Learnable from Rewards?

AI 导读

研究探讨预训练与中训练如何为奖励适配提供信息与计算,用 Qwen2.5 检查点验证分工。八个世界中,接受正确来源与首操作监督的 Sequential 模型成功率达 82.61%,私有随机来源对照组仅 44.15%;独立八世界确认实验为 75.32% 对 49.86%。GSM8K 与 HotpotQA 区分了奖励进入时的准确率、后续增益与最终表现。

正文

View PDF HTML (experimental)

Abstract:A reward can identify a correct answer while leaving the computation needed for new inputs undetermined. We study how pretraining and midtraining supply the information and computation that make reward adaptation effective. In sequential state computation and contextual memory, we characterize mechanisms that agree on every training reward yet demand different held-out answers. Task-independent source observations resolve this ambiguity. We construct finite sampled Adam paths from specified random initializations through source prediction and reward adaptation in the same parameters, proving how prediction acquires execution or retrieval and rewards learn their task-specific use. Experiments with pretrained Qwen2.5 checkpoints test this division of labor. Across eight worlds, Sequential models trained with correct source and first-operation supervision reach 82.61% success, versus 44.15% for a private-random source control. Memory replay preserves retrieval during reward adaptation, and an independent eight-world confirmation achieves 75.32% task success versus 49.86% after matched alternative-retrieval training. GSM8K and HotpotQA separate accuracy at reward entry, subsequent gain and final performance. Together, these results connect information acquisition, executable computation and reward-guided task learning.
Comments: 160 pages, 26 figures, 32 tables
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2609.38446 [cs.LG]
  (or arXiv:2609.38446v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.38446

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Chiwun Yang [view email]
[v1] Tue, 29 Sep 2026 19:39:41 UTC (569 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org