arXiv:cs.LG(机器学习,全量分类)· Michael Sullivan, Alexander Koller·· 14 小时前AI 评分54
arXiv 论文分析 RLVR 后训练中的语言漂移:理论上无上界且难以约束
On Language Drift during RLVR Post-Training
AI 导读
Michael Sullivan 与 Alexander Koller 在 arXiv 论文(arXiv:2610.02015)中研究 LLM 推理模型在 RLVR 后训练中出现的语言漂移现象。
正文
Abstract:Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)---have enabled them to accomplish impressively complex tasks. However, in parallel with their rising capabilities, LLMs have increasingly displayed signs of language drift in their chains of thought (CoTs): unusual, non-standard, and seemingly nonsensical language use. Although it is well-documented---and can potentially impair CoT monitorability---the causes of language drift are thus far poorly understood. In this paper, we identify the conditions under which language drift occurs: we prove theoretically that RLVR optimization pressure permits unbounded language drift, while supervised fine-tuning does not. We then show empirically that language drift specifically arises during RLVR on novel reasoning tasks---i.e. when the target behavior cannot be drawn out of the base model. Finally, we prove that it is not possible to constrain language drift without constraining expected reward, suggesting that CoT monitorability cannot be improved without harming performance during RLVR post-training at the frontier.
| Comments: | 22 pages; 15 figures; 4 tables |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.02015 [cs.LG] |
| (or arXiv:2610.02015v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02015 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Michael Sullivan [view email]
[v1]
Thu, 1 Oct 2026 16:39:48 UTC (899 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org