跳到正文
原文
HuggingFace Daily Papers(社区热门论文)·· 12 小时前AI 评分34

Neighborhood OPSD:邻域在线策略自蒸馏提升数学推理

Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation

AI 导读

研究者提出 Neighborhood OPSD(N-OPSD),通过局部参数扰动构建冻结专家池,在相同参考解下提供互补的参考对齐修正,将邻域监督用于学生访问状态。

正文

Published on Sep 30

·

Submitted by

dingyi

on Oct 2

Authors:

,

,

,

,

,

,

,

Abstract

On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher that sees a reference solution and supervises student-sampled prefixes. Standard OPSD uses one fixed parameter setting at every state, but nearby settings may offer additional supervision. We find that local parameter perturbations reveal complementary reference-aligned corrections under the same reference context. Different experts supply these corrections at different reference positions. Their pool covers more such positions than the unperturbed privileged teacher. We introduce Neighborhood OPSD (N-OPSD) to turn these corrections into supervision at student-visited states. Offline, greedy selection builds a compact pool of frozen experts by rewarding filtered reference-token gains beyond the pool's current best at each position. The highest-peak expert need not provide the best training target. Online routing therefore separates the anchor direction from its level of support. MaxPeak selects the anchor token, and quantile selection chooses among experts whose top token matches it. The student learns from the chosen expert's full next-token distribution through the clipped forward-KL objective inherited from OPSD. We evaluate on AIME 2024, AIME 2025, and HMMT February 2025. Across three independent runs per method, Neighborhood OPSD improves the three-benchmark Average@12 over OPSD by 2.75, 1.67, and 1.94 points on Qwen3-1.7B, 4B, and 8B, respectively. Student-prefix continuations support using the pool beyond the reference trajectories used for selection. Matched ablations support filtered reference-token gains as a selection criterion. Accounting for overlap within the pool and routing by state further improve student accuracy. Inference uses only the distilled student.

View arXiv page View PDF Project page Add to collection

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.39687 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.39687 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.39687 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co