arXiv:cs.LG(机器学习,全量分类)· Junxian Li, Ruixuan Yang, Tianao Zhang, Tiange Xu, Weisheng Dong, Yulun Zhang·· 1 天前AI 评分45
LT-OPD:面向极端视觉 token 压缩的在线策略自蒸馏
Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction
AI 导读
研究者提出 LT-OPD 训练框架,让仅保留少量视觉 token 的学生模型沿自身生成轨迹接受冻结全 token 教师模型的分布监督,并引入预算级课程逐步降低 token 预算。
正文
Abstract:Visual token reduction is an effective way to accelerate multimodal large language models (MLLMs), but performance deteriorates rapidly under extremely low token budgets. Existing work has explored both visual-token selection and training-based adaptation to reduced visual inputs. We take a step further by asking how a heavily compressed MLLM should learn from the states induced by its own generations. This setting naturally calls for on-policy self-distillation: a heavily compressed model is supervised on the states induced by its own generations, while its full-token counterpart serves as an information-rich teacher. Based on this insight, we propose LT-OPD, a training framework for extreme visual-token reduction. The student rolls out responses with only a small fraction of visual tokens, and a frozen full-token copy of the same MLLM provides distributional supervision along these student-generated trajectories. To stabilize on-policy learning when visual evidence is severely limited, we further introduce a budget-level curriculum that progressively decreases the token budget during training. Across nine benchmarks on Qwen3.5-4B, LT-OPD raises average retained performance under 5% visual-token retention from 68.6% to 82.3%, outperforming training-free, training-based, and reinforcement-learning baselines at the same budget. The gains transfer consistently to Qwen3.5-9B, GLM-4.6V-9B, and LLaVA-OV-1.5-4B. LT-OPD also reduces KV-cache usage by 85.2% and prefill FLOPs by 85.4% without additional inference overhead, demonstrating that on-policy learning can substantially recover capabilities lost to extreme visual-token reduction.
| Comments: | Code is at this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.32353 [cs.CV] |
| (or arXiv:2609.32353v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.32353 arXiv-issued DOI via DataCite |
Submission history
From: Junxian Li [view email]
[v1]
Sat, 26 Sep 2026 08:16:07 UTC (2,817 KB)
[v2]
Thu, 1 Oct 2026 14:20:19 UTC (2,794 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org