跳到正文
arXiv:cs.CL· Xiaoshu Chen, Sihang Zhou, Ke Liang, Xinwang Liu·· 4 小时前AI 评分38

SeOPD:通过自生成思维链进行在线策略蒸馏的自进化 LLM

SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought

AI 导读

SeOPD 让 LLM 无需外部特权信息即可自我提升:用深度思考模式生成 CoT,将其作为特权信息对非思考模式的回答做 token 级监督,使推理中产生的新信息内化到共享参数中,从而同步提升两种模式能力。该方法在不依赖人工标注或外部环境反馈的前提下实现了自我改进,在多种 LLM 和任务上的实验验证了其有效性。

正文

View PDF HTML (experimental)

Abstract:Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or feedback from external environments. However, obtaining accurate annotations and constructing sophisticated environments often require substantial human effort and computation, limiting the scalability of OPSD. While a few recent studies have explored self-improvement without external PI, the resulting gains remain limited. In this work, we explore whether LLMs can achieve comparable self-improvement without external PI. Our key observation is that a single LLM can support multiple reasoning modes, such as deep-thinking and non-thinking modes, with deep thinking generating additional information during reasoning. Based on this observation, we propose Self-Evolving Online Policy Distillation (SeOPD), which enables LLMs to distill and internalize information generated by their own chain of thought (CoT). Specifically, it (1) generates CoT with the deep-thinking mode, (2) produces responses with the non-thinking mode, and (3) uses the generated CoT as PI to provide token-level supervision for the non-thinking response, allowing new information inferred during reasoning to guide the non-thinking mode and be internalized into the shared model parameters, thereby improving both non-thinking and deep-thinking capabilities. Extensive experiments across LLMs and tasks demonstrate the effectiveness of SeOPD.
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2609.33181 [cs.AI]
  (or arXiv:2609.33181v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.33181

arXiv-issued DOI via DataCite

Submission history

From: Xiaoshu Chen [view email]
[v1] Sun, 27 Sep 2026 04:00:06 UTC (562 KB)
[v2] Wed, 7 Oct 2026 12:52:13 UTC (557 KB)

来源:arXiv:cs.CL · arxiv.org