arXiv:cs.LG(机器学习,全量分类)· Minho Jeong, Dooho Lee, Jinmo Lee, Jaemin Yoo·· 14 小时前AI 评分44
将表格基础模型蒸馏为高效预测器
Distillation of Tabular Foundation Models into Efficient Predictors
AI 导读
研究提出一套表格基础模型(TFM)知识蒸馏方案:以完整标注训练集作为教师上下文,仅用教师对真实与合成查询的预测训练轻量学生模型。在 TabArena 上,蒸馏学生比有监督调优集成模型高出 57-98 Elo 分;原样应用于 TALENT 时,在 300 个数据集中的 236-258 个上优于默认学生,主误差中位数降低 4.0-6.4%,推理速度中位数提升 3.0-21.6 倍。
正文
Abstract:Tabular foundation models (TFMs) achieve strong predictive performance through in-context learning, yet repeatedly conditioning on labeled data makes inference expensive. Knowledge distillation can reduce this cost by transferring their predictive ability to lightweight, dataset-specific students. However, the dependence of TFM predictions on both a labeled context and a query introduces two design questions: how to construct teacher supervision and whether expanding query coverage improves distillation. We examine these questions across two TFMs and both neural and tree-based students, and derive an effective distillation recipe. The recipe uses the full labeled training set as teacher context and trains students solely on teacher predictions for observed and synthetic queries. On TabArena, the resulting students outperform their supervised trained tuned-and-ensembled counterparts by 57-98 Elo points. Applied unchanged to TALENT, the same recipe improves matched default students on 236-258 of 300 datasets and reduces median primary error by 4.0-6.4%. The distilled students also achieve median inference speedups of 3.0-21.6 times over their teachers, offering a practical trade-off between predictive performance and repeated inference cost. Code is available at this https URL .
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.01435 [cs.LG] |
| (or arXiv:2610.01435v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01435 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Minho Jeong [view email]
[v1]
Thu, 1 Oct 2026 10:31:22 UTC (291 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org