arXiv:cs.LG(机器学习,全量分类)· Jaeyun Shin, Hangeol Chang, Jong Chul Ye·· 13 小时前AI 评分41
LEGO-OPD:面向多模态同策略蒸馏的因子化教师组合方法
LEGO-OPD: Factorized Teacher Composition for Multimodal On-Policy Distillation
AI 导读
LEGO-OPD 通过将语言专家和视觉定位专家的预测因子化组合为单一教师分布,实现多模态同策略蒸馏中语言推理与视觉定位的独立控制。该方法引入自适应校准,以定位专家的图像诱导预测偏移作为前缀相关参考,避免视觉监督不足或过度。基于 Qwen3 的实验显示,LEGO-OPD 在多模态和纯文本推理任务上均优于单教师和多教师 OPD 基线。
正文
Abstract:Multimodal on-policy distillation (OPD) aims to improve visual grounding while preserving the strong reasoning capabilities of language models. Recent multi-teacher approaches combine LLM and VLM teachers to provide complementary supervision. However, directly using a VLM's full predictive distribution entangles its visual grounding signal with its own language prior, preventing the grounding information from being transferred independently. Conversely, increasing the strength of visual supervision can improve perception but may overemphasize visual evidence and degrade language reasoning. To address this trade-off, we introduce LEGO-OPD, which selectively composes factors from a Language Expert and a Grounding expert into One teacher distribution for multimodal OPD. Under a generalized Bayesian formulation, the language expert provides a prior over candidate tokens, while the grounding expert contributes a visual likelihood that updates this prior, rather than transferring its complete predictive distribution. This factorized composition allows language reasoning and visual grounding to be controlled independently. We further introduce adaptive calibration to determine how strongly the visual likelihood should update the language prior at each decoding prefix. Specifically, LEGO-OPD uses the grounding expert's image-induced prediction shift as a prefix-dependent reference, preventing both insufficient and excessive visual supervision. Experiments with Qwen3 models show that LEGO-OPD consistently outperforms the evaluated single- and multi-teacher OPD baselines on both multimodal and text-only reasoning tasks. Moreover, it improves the initial student's visual perception while preserving text-only reasoning.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00333 [cs.CV] |
| (or arXiv:2610.00333v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00333 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jong Chul Ye [view email]
[v1]
Tue, 29 Sep 2026 13:57:51 UTC (292 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org