arXiv:cs.AI· Nan Qiao, Yebin Yang, Weinong Wang, Shuning Wang, Shangpin Peng, Fengyuan Lu, Xinming Wang, Zhehan Kan, Ruixu Zhang, Songyang Zhang, Sheng Yue, Yonglong Tian, Ju Ren·· 4 小时前AI 评分37
T⁵:面向强化中期训练中 Token 级思维的双评论家训练方法
$T^5$: Twin-Critic Training for Token-Level Thoughts in Reinforcement Mid-Training
AI 导读
研究者提出双评论家方法 T⁵,通过单条生成轨迹校准 Token 级优势值,用于强化中期训练中的 Token 级思维学习。该方法在预热与留出验证后给出两个优势估计,并以动作相关权重经条件矩鞍点目标融合,使各前缀平均优势趋近零,同时用信号保留约束防止学习信号被抹除。实验显示,相比当前最优的无评论家方法,T⁵ 平均基准性能提升 7.8%,平均训练步耗时最多降低 63.4%。
正文
Authors:Nan Qiao, Yebin Yang, Weinong Wang, Shuning Wang, Shangpin Peng, Fengyuan Lu, Xinming Wang, Zhehan Kan, Ruixu Zhang, Songyang Zhang, Sheng Yue, Yonglong Tian, Ju Ren
Abstract:Reinforcement mid-training lets language models learn internal thoughts from unlabeled text, but efficient token-level credit assignment remains challenging. Existing group-relative methods require costly repeated generation. Learned critics offer single-rollout feedback, but accurate return prediction alone does not ensure reliable policy updates. Our analysis shows how training--inference mismatch and PPO clipping prevent a common offset in advantage estimates from cancelling out, introducing additional update drift. We propose \tfour{}, a twin-critic method that calibrates token-level advantages from a single generated trajectory. After warmup and held-out qualification, the critics provide two advantage estimates, combined using action-dependent weights learned through a conditional-moment saddle-point objective. This objective brings the average advantage at each prefix toward zero, while a signal-retention constraint prevents the correction from erasing the learning signal. Sharing information across text positions avoids repeated sampling of each prefix. Theoretically, we characterize optimal mixing under the signal-retention constraint and establish an upper bound on residual mean-induced drift. Experiments show that, compared with the state-of-the-art critic-free method, \tfour{} improves mean benchmark performance by 7.8\% and reduces mean training-step time by up to 63.4\%.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.32791 [cs.AI] |
| (or arXiv:2609.32791v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.32791 arXiv-issued DOI via DataCite |
Submission history
From: Nan Qiao [view email]
[v1]
Sat, 26 Sep 2026 17:07:57 UTC (508 KB)
[v2]
Fri, 2 Oct 2026 16:25:36 UTC (513 KB)
来源:arXiv:cs.AI · arxiv.org