跳到正文
arXiv:cs.CL· Houcheng Jiang, Mao Zheng, Mingyang Song, Qiyong Zhong, Jie Sun, Tianyu Zhang, Junfeng Fang·· 3 小时前

ReCal:面向同策略蒸馏恢复的结构化剪枝校准方法

ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery

AI 导读

研究者提出 RECAL(Recovery-Aware Calibration),一种即插即用的剪枝前校准方法,通过未剪枝教师模型与剪枝探针之间的前向 KL 识别被剪枝破坏的教师支持预测,并重新加权校准统计量,引导现有剪枝准则保留这些预测。在多种模型和剪枝方法上,RECAL 使同策略蒸馏(OPD)后的数学推理表现持续提升,AIME 最高提升 16.7 个百分点,多数代码生成对比也有改善。

正文

View PDF HTML (experimental)

Abstract:Structured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery. Because OPD relies on student-generated trajectories, pruning damage that persists after offline distillation can limit its effectiveness. We propose RECAL, Recovery-Aware Calibration, a simple plug-and-play approach that improves OPD recovery by adjusting calibration before pruning. RECAL uses forward KL between an unpruned teacher and a pruned probe to identify teacher-supported predictions disrupted by pruning, then reweights calibration statistics to guide existing pruning criteria toward preserving these predictions. Across multiple models and pruning methods, RECAL consistently improves mathematical reasoning after OPD, achieving gains of up to 16.7 percentage points on AIME, alongside improvements in most code-generation comparisons. Further analysis shows that RECAL reduces residual damage at heavily affected tokens and establishes performance advantages that persist through recovery. These results demonstrate the value of recovery-aware calibration for improving on-policy distillation recovery of pruned reasoning models.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.11332 [cs.CL]
  (or arXiv:2610.11332v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.11332

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Houcheng Jiang [view email]
[v1] Thu, 8 Oct 2026 06:28:38 UTC (309 KB)

来源:arXiv:cs.CL · arxiv.org