跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Matthieu Zimmer, Xiaotong Ji, Tu Nguyen, Haitham Bou-Ammar·· 14 小时前AI 评分37

用最坏情况约束强化学习蒸馏 LLM 推理能力

The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning

AI 导读

研究者将 LLM 推理蒸馏建模为约束强化学习问题,在轨迹每个前缀上对教师对数似然施加最坏情况约束,从而最大化任务奖励。该方法推导出无需状态增广的约束 MDP,其策略梯度可分解为单步与长期项,并在惩罚极限下几乎必然满足最坏情况约束。在数学推理与代码生成任务上,它同时匹配纯 RL 的最终答案正确率并大幅减少教师约束违反,取得最高严格推理成功率。

正文

View PDF HTML (experimental)

Abstract:Distilling the reasoning capabilities of large language models (LLMs) into smaller students is a central challenge for efficient deployment. Current approaches face a fundamental tension: optimizing purely for verifiable task rewards (e.g., via GRPO) leads to reward hacking, where students arrive at correct final answers through flawed intermediate logic, while regularizing with soft divergence penalties against a teacher (e.g., KL-based distillation) dilutes task performance and, critically, allows the student to compensate for severe logical violations at one step with high teacher agreement at others. We argue that this averaging is fundamentally misaligned with the nature of reasoning: a chain-of-thought is only as valid as its weakest link. Motivated by this observation, we formulate reasoning distillation as a constrained reinforcement learning problem in which the task reward is maximized subject to a worst-case constraint on the teacher log-likelihood along every prefix of the trajectory. To avoid the prohibitive cost of dual Lagrangian solvers and the test-time teacher dependence of state-augmented methods such as Saute, we derive an unaugmented constrained MDP whose reward transformation preserves the hard-constraint semantics, admits a low-variance policy gradient decomposition into single-step and long-term terms, and provably satisfies the worst-case constraint almost surely in the penalty limit. Through extensive experiments on mathematical reasoning and code generation tasks, we demonstrate that our method significantly expands the accuracy-fidelity Pareto front. By matching the high Final Answer Correctness of pure RL and drastically reducing teacher constraint violations, we ultimately achieve the highest rigorous Reasoning Success Rate across all evaluated settings.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.00332 [cs.LG]
  (or arXiv:2610.00332v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00332

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xiaotong Ji [view email]
[v1] Tue, 29 Sep 2026 13:51:52 UTC (155 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org