arXiv:cs.AI· Haohui Wang, Jiahao Xu, Wangzhi Zhan, Tong Zeng, Dongqi Fu, Hong Li, Swastik Roy, Naren Ramakrishnan, Chris North, Jian Kang, Yujun Yan, Dawei Zhou·· 3 小时前
长尾分布下的监督微调:PASS 方法如何克服先验壁垒
Overcoming Prior Barriers: Supervised Fine-Tuning under Long-Tail Distribution
AI 导读
研究者提出"先验壁垒"概念,量化预训练模型对竞争概念的支撑强度,发现其呈长尾分布——高频概念壁垒低、尾部概念壁垒高,SFT 需更多指令才能克服。据此提出自适应指令选择方法 PASS,联合考虑哪些指令能提供有效证据与何处需要额外监督。实验表明 PASS 在四组骨干-预算设置下持续优于七种 SOTA 指令选择方法,消融研究显示其自适应分配优于均匀分配。
正文
Authors:Haohui Wang, Jiahao Xu, Wangzhi Zhan, Tong Zeng, Dongqi Fu, Hong Li, Swastik Roy, Naren Ramakrishnan, Chris North, Jian Kang, Yujun Yan, Dawei Zhou
Abstract:Supervised fine-tuning (SFT) adapts pretrained large language models (LLMs) to downstream tasks, but the required concepts can receive substantially different levels of pretrained support. Frequent concepts are more likely to be well learned, whereas rare concepts may remain weakly represented. We introduce a novel notion named prior barrier to quantify how strongly the pretrained model supports competing concepts over the target concept. We observe that prior barriers follow a long-tail distribution, placing head and tail concepts at different starting points for SFT: head concepts face lower prior barriers, whereas tail concepts require additional instructions to overcome their higher prior barriers. Our theoretical analysis further derives a predictive risk bound for SFT under long-tail prior barriers, explicitly characterizing how the prior barrier and accumulated SFT evidence jointly determine predictive performance. Motivated by this prior barrier-dependent demand, we propose PASS, an adaptive SFT instruction selection method that constructs reference-derived concepts and estimates the distinguishing evidence provided by each instruction, and adaptively allocates the selection budget toward concepts that remain insufficiently covered under the current selection. In this way, PASS jointly considers which instructions can provide useful evidence and where additional supervision is needed under a limited budget. Experiments show that our method consistently outperforms seven state-of-the-art instruction selection methods on four backbone-budget settings. An ablation study further shows that PASS's adaptive allocation consistently improves over uniform allocation.
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.12345 [cs.AI] |
| (or arXiv:2610.12345v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12345 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Haohui Wang [view email]
[v1]
Thu, 8 Oct 2026 17:16:45 UTC (1,697 KB)
来源:arXiv:cs.AI · arxiv.org