arXiv:cs.LG· Haifeng Wu, Srinivasan Manoharan, Jian Wan, Fangbo Tu, Junhua Zhao, Xin Chen·· 3 小时前AI 评分31
学生引导的教师蒸馏实现高效 LLM 任务路由:与 Jev 式 System-1 分类器的对比定位
Student-Guided Teacher Distillation for Efficient LLM Task Routing: Positioning Against Jev-Style System-1 Classifiers
AI 导读
研究提出一种学生引导的教师蒸馏流程,用于 60 类 LLM 任务分类:ModernBERT 学生模型一次前向预测全类别分布并召回 top-k 候选,再由 DeBERTa-v3 零样本 NLI 教师仅对候选重排,教师标签迭代改进学生。
正文
Abstract:Zero-shot classifiers are useful for routing user requests to specialized LLM tasks, but scoring every request against a large candidate set is expensive: a zero-shot NLI classifier must evaluate one premise-hypothesis pair per label, so cost scales linearly with taxonomy size. We study a student-guided teacher distillation pipeline for a fixed taxonomy of 60 LLM task categories: a compact ModernBERT classifier predicts the full category distribution in one forward pass and retrieves a small top-k candidate set, and a larger DeBERTa-v3 zero-shot NLI classifier reranks only those candidates rather than all 60 labels; the resulting teacher labels iteratively improve the student, which produces sharper candidates for the next round. Unlike generic embedding retrieval or clustering-derived shortlists used in extreme multi-label classification, our candidate generator is trained end-to-end on the target taxonomy and is the same model serving production traffic, distinguishing it from LLM-routing work that routes between candidate models, and from concurrent System-1 encoder-classifier proposals (e.g. TypeSafe AI's Jev and the open-source Laya project) whose training methodology is undocumented or RL-based. Our best student checkpoint reaches 77.5% teacher agreement on a 200-example evaluation set, and preliminary coverage measurements show Coverage@16 of 91-100%, suggesting top-k sets retain most of the teacher's decision-relevant information. We further show truncated top-k teacher scores should not be treated as full 60-class soft targets for KL distillation: zeroing untruncated classes destroys the dark knowledge soft-label distillation depends on, introducing systematic bias rather than a harmless sparse approximation. A complete evaluation, including coverage at multiple k on a held-out set, an embedding-retrieval baseline, and a larger human-reviewed test set, remains in progress.
| Comments: | 13 pages, 2 figures, 1 table |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| ACM classes: | I.2.7; I.2.6 |
| Cite as: | arXiv:2610.02516 [cs.LG] |
| (or arXiv:2610.02516v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02516 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Haifeng Wu [view email]
[v1]
Thu, 1 Oct 2026 21:38:11 UTC (17 KB)
来源:arXiv:cs.LG · arxiv.org