跳到正文
arXiv:cs.AI· Nan Qiao, Yebin Yang, Weinong Wang, Shuning Wang, Shangpin Peng, Fengyuan Lu, Xinming Wang, Zhehan Kan, Ruixu Zhang, Songyang Zhang, Sheng Yue, Yonglong Tian, Ju Ren·· 4 小时前AI 评分37

T⁵:面向强化中期训练中 Token 级思维的双评论家训练方法

$T^5$: Twin-Critic Training for Token-Level Thoughts in Reinforcement Mid-Training

AI 导读

研究者提出双评论家方法 T⁵,通过单条生成轨迹校准 Token 级优势值,用于强化中期训练中的 Token 级思维学习。该方法在预热与留出验证后给出两个优势估计,并以动作相关权重经条件矩鞍点目标融合,使各前缀平均优势趋近零,同时用信号保留约束防止学习信号被抹除。实验显示,相比当前最优的无评论家方法,T⁵ 平均基准性能提升 7.8%,平均训练步耗时最多降低 63.4%。

正文

Authors:Nan Qiao, Yebin Yang, Weinong Wang, Shuning Wang, Shangpin Peng, Fengyuan Lu, Xinming Wang, Zhehan Kan, Ruixu Zhang, Songyang Zhang, Sheng Yue, Yonglong Tian, Ju Ren

View PDF HTML (experimental)

Abstract:Reinforcement mid-training lets language models learn internal thoughts from unlabeled text, but efficient token-level credit assignment remains challenging. Existing group-relative methods require costly repeated generation. Learned critics offer single-rollout feedback, but accurate return prediction alone does not ensure reliable policy updates. Our analysis shows how training--inference mismatch and PPO clipping prevent a common offset in advantage estimates from cancelling out, introducing additional update drift. We propose \tfour{}, a twin-critic method that calibrates token-level advantages from a single generated trajectory. After warmup and held-out qualification, the critics provide two advantage estimates, combined using action-dependent weights learned through a conditional-moment saddle-point objective. This objective brings the average advantage at each prefix toward zero, while a signal-retention constraint prevents the correction from erasing the learning signal. Sharing information across text positions avoids repeated sampling of each prefix. Theoretically, we characterize optimal mixing under the signal-retention constraint and establish an upper bound on residual mean-induced drift. Experiments show that, compared with the state-of-the-art critic-free method, \tfour{} improves mean benchmark performance by 7.8\% and reduces mean training-step time by up to 63.4\%.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.32791 [cs.AI]
  (or arXiv:2609.32791v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.32791

arXiv-issued DOI via DataCite

Submission history

From: Nan Qiao [view email]
[v1] Sat, 26 Sep 2026 17:07:57 UTC (508 KB)
[v2] Fri, 2 Oct 2026 16:25:36 UTC (513 KB)

来源:arXiv:cs.AI · arxiv.org