LoopFM:从基础模型历史表示中学习,突破推荐系统知识蒸馏瓶颈
LoopFM: Learning frOm HistOrical RePresentations of Foundation Model for Recommendation
LoopFM 将基础模型(FM)的中间嵌入向量结构化为下游垂直模型的输入特征,无需实时 FM 推理即可实现高带宽知识迁移。在三个公开基准上 AUC 提升显著,如 TaobaoAd 上提升 6%+;工业级规模(数十亿样本、万亿参数 FM)下,其知识蒸馏比率约为 KD 的两倍,首次上线后转化率提升 +0.5%。
Authors:Hua Zheng, Shali Jiang, Boyang Liu, Laming Chen, Kenny Lov, Chuanqi Xu, Lisang Ding, Qinghai Zhou, Can Cui, Xiaolong Liu, Xiaoyi Liu, Yasmine Badr, Xin Xu, Mingfu Liang, Jiyan Yang, Ellie Dingqiao Wen, Gerard Jonathan Mugisha Akkerhuis, Jason Rudy, Xi Liu, Chenxiao Guan, Rong Jin, Ruichao Qiu, Xian Chen, Zhehui Zhou, Ping Chen, Rui Yang, Haicheng Chen, Meet Raval, Song Zhou, Dharak Kharod, Shuyu Xu, Xingyuan Wang, Liang Tao, Qiang Jin, Qiao Yang, Wankun Zhu, Qin Huang, Yuzhen Huang, Darren Liu, Parish Aggarwal, Hui Zhou, Erzhuo Wang, Shuo Chang, Xiaorui Gan, Wenlin Chen, Santanu Kolay, Huayu Li
Abstract:Knowledge distillation (KD) transfers a single scalar prediction from a large foundation model (FM) to compact vertical models (VMs), suffering from diminishing transfer ratio -- the fraction of FM improvement captured by the VM -- as a single scalar cannot convey the rich intermediate knowledge that larger FMs learn. To address this bottleneck, we propose LoopFM (Learning frOm HistOrical RePresentations of FM), a framework that opens a high-bandwidth transfer channel by structuring FM intermediate embeddings as input features (e.g., user history sequence) for downstream VMs, without requiring real-time FM inference at serving and architectural coupling between FM and VM. We provide a theoretical framework for LoopFM with a gain decomposition and transfer-ratio analysis. On three public benchmarks, LoopFM demonstrates strong AUC improvements (e.g., 6%+ on TaobaoAd) and complementary knowledge transfer capability with KD. On industrial-scale systems (billions of examples, trillion-parameter FMs), LoopFM approximately doubles the knowledge transfer ratio on top of KD, delivering a +0.5% conversion improvement in the first half after its initial launch, and +1.03% and +1.22% conversion improvement from two individual launches in the subsequent half. Through systematic experiments, LoopFM demonstrates a scaling law in sequence length, embedding dimension, and upstream FM size.
| Comments: | Hua Zheng, Shali Jiang, Boyang Liu contributed equally to this work |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR) |
| Cite as: | arXiv:2605.29280 [cs.LG] |
| (or arXiv:2605.29280v3 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.29280 arXiv-issued DOI via DataCite |
Submission history
From: Hua Zheng [view email]
[v1]
Thu, 28 May 2026 02:59:46 UTC (1,450 KB)
[v2]
Tue, 2 Jun 2026 18:06:36 UTC (1,214 KB)
[v3]
Wed, 7 Oct 2026 03:55:40 UTC (1,229 KB)
来源:arXiv:cs.LG · arxiv.org