更小的模型,更好的拒绝样本:偏好蒸馏的规模规律
Smaller Models, Better Rejects: Preference Distillation Scaling
偏好蒸馏研究发现在 7B 到 72B 的学生模型上,用更小的冻结模型生成拒绝样本比自生成拒绝样本推理成本更低、训练效果更好,在代码生成和数学推理任务上均成立。研究推导出 DPO 的有限时域效用上界,并据此提出三种干预:混合更小模型与学生规模模型的拒绝样本、将拒绝样本重新分配给其他提示词、优先选择参考策略下似然更低的候选样本。
Published on Sep 30
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither assumption holds: across students from 7B to 72B, smaller frozen models generate rejects with less inference compute yet train stronger students than self-generated rejects, before and after sequence-level knowledge distillation, on code generation and mathematical reasoning. To explain this result, we derive a finite-horizon utility bound for Direct Preference Optimization in a linearized feature model. The bound characterizes favorable reject distributions and motivates three interventions. First, mixing rejects from smaller and student-scale models improves performance as the smaller model's share increases. Second, reassigning rejects to other prompts and shuffling their code tokens still outperform length-matched gibberish, showing that task structure contributes to reject utility. Third, selecting candidates with lower likelihood under the reference policy improves net transfer when higher-likelihood candidates provide less useful contrast. Lower-likelihood selections outperform higher-likelihood ones for every source. These results suggest that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide them at low cost.
View arXiv page View PDF Add to collection
Get this paper in your agent:
hf papers read 2609.38987
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this paper 0
No model linking this paper
Cite arxiv.org/abs/2609.38987 in a model README.md to link it from this page.
Datasets citing this paper 0
No dataset linking this paper
Cite arxiv.org/abs/2609.38987 in a dataset README.md to link it from this page.
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2609.38987 in a Space README.md to link it from this page.
Collections including this paper 0
No Collection including this paper
Add this paper to a collection to link it from this page.
来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co