跳到正文
arXiv:cs.LG· Haifeng Wu, Srinivasan Manoharan, Jian Wan, Fangbo Tu, Junhua Zhao, Xin Chen·· 3 小时前AI 评分31

学生引导的教师蒸馏实现高效 LLM 任务路由:与 Jev 式 System-1 分类器的对比定位

Student-Guided Teacher Distillation for Efficient LLM Task Routing: Positioning Against Jev-Style System-1 Classifiers

AI 导读

研究提出一种学生引导的教师蒸馏流程,用于 60 类 LLM 任务分类:ModernBERT 学生模型一次前向预测全类别分布并召回 top-k 候选,再由 DeBERTa-v3 零样本 NLI 教师仅对候选重排,教师标签迭代改进学生。

正文

View PDF HTML (experimental)

Abstract:Zero-shot classifiers are useful for routing user requests to specialized LLM tasks, but scoring every request against a large candidate set is expensive: a zero-shot NLI classifier must evaluate one premise-hypothesis pair per label, so cost scales linearly with taxonomy size. We study a student-guided teacher distillation pipeline for a fixed taxonomy of 60 LLM task categories: a compact ModernBERT classifier predicts the full category distribution in one forward pass and retrieves a small top-k candidate set, and a larger DeBERTa-v3 zero-shot NLI classifier reranks only those candidates rather than all 60 labels; the resulting teacher labels iteratively improve the student, which produces sharper candidates for the next round. Unlike generic embedding retrieval or clustering-derived shortlists used in extreme multi-label classification, our candidate generator is trained end-to-end on the target taxonomy and is the same model serving production traffic, distinguishing it from LLM-routing work that routes between candidate models, and from concurrent System-1 encoder-classifier proposals (e.g. TypeSafe AI's Jev and the open-source Laya project) whose training methodology is undocumented or RL-based. Our best student checkpoint reaches 77.5% teacher agreement on a 200-example evaluation set, and preliminary coverage measurements show Coverage@16 of 91-100%, suggesting top-k sets retain most of the teacher's decision-relevant information. We further show truncated top-k teacher scores should not be treated as full 60-class soft targets for KL distillation: zeroing untruncated classes destroys the dark knowledge soft-label distillation depends on, introducing systematic bias rather than a harmless sparse approximation. A complete evaluation, including coverage at multiple k on a held-out set, an embedding-retrieval baseline, and a larger human-reviewed test set, remains in progress.
Comments: 13 pages, 2 figures, 1 table
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
ACM classes: I.2.7; I.2.6
Cite as: arXiv:2610.02516 [cs.LG]
  (or arXiv:2610.02516v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02516

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Haifeng Wu [view email]
[v1] Thu, 1 Oct 2026 21:38:11 UTC (17 KB)

来源:arXiv:cs.LG · arxiv.org