跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Junxian Li, Ruixuan Yang, Tianao Zhang, Tiange Xu, Weisheng Dong, Yulun Zhang·· 1 天前AI 评分45

LT-OPD:面向极端视觉 token 压缩的在线策略自蒸馏

Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction

AI 导读

研究者提出 LT-OPD 训练框架,让仅保留少量视觉 token 的学生模型沿自身生成轨迹接受冻结全 token 教师模型的分布监督,并引入预算级课程逐步降低 token 预算。

正文

View PDF HTML (experimental)

Abstract:Visual token reduction is an effective way to accelerate multimodal large language models (MLLMs), but performance deteriorates rapidly under extremely low token budgets. Existing work has explored both visual-token selection and training-based adaptation to reduced visual inputs. We take a step further by asking how a heavily compressed MLLM should learn from the states induced by its own generations. This setting naturally calls for on-policy self-distillation: a heavily compressed model is supervised on the states induced by its own generations, while its full-token counterpart serves as an information-rich teacher. Based on this insight, we propose LT-OPD, a training framework for extreme visual-token reduction. The student rolls out responses with only a small fraction of visual tokens, and a frozen full-token copy of the same MLLM provides distributional supervision along these student-generated trajectories. To stabilize on-policy learning when visual evidence is severely limited, we further introduce a budget-level curriculum that progressively decreases the token budget during training. Across nine benchmarks on Qwen3.5-4B, LT-OPD raises average retained performance under 5% visual-token retention from 68.6% to 82.3%, outperforming training-free, training-based, and reinforcement-learning baselines at the same budget. The gains transfer consistently to Qwen3.5-9B, GLM-4.6V-9B, and LLaVA-OV-1.5-4B. LT-OPD also reduces KV-cache usage by 85.2% and prefill FLOPs by 85.4% without additional inference overhead, demonstrating that on-policy learning can substantially recover capabilities lost to extreme visual-token reduction.
Comments: Code is at this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2609.32353 [cs.CV]
  (or arXiv:2609.32353v2 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.32353

arXiv-issued DOI via DataCite

Submission history

From: Junxian Li [view email]
[v1] Sat, 26 Sep 2026 08:16:07 UTC (2,817 KB)
[v2] Thu, 1 Oct 2026 14:20:19 UTC (2,794 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org