arXiv:cs.CL· Ziran Li, Qiang Wang, Zhengyu Chen, Shanglin Lei, Borun Chen, Jingang Wang, Xunliang Cai·· 4 小时前
从文本到视觉能迁移什么?VLM 的能力缩放定律与迁移动态
What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs
AI 导读
研究者提出 Capability-Driven Multimodal Scaling Law,首个跨模型家族框架,可从 LLM 文本基准的 PCA 能力分数 S 预测 VLM 基准准确率。
正文
Abstract:Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains fundamentally unprincipled: compute-based scaling laws fail to generalize across model families, and no framework exists for directly predicting VLM performance before training begins. We propose the Capability-Driven Multimodal Scaling Law, the first cross-family framework that predicts VLM benchmark accuracy from directly observable textual capability. Given a low-dimensional capability score $S$ extracted from LLM textual benchmarks via PCA, we model VLM performance as a function of $S$, with a per-backbone transfer rate and an absorption rate that quantifies data-scaling efficiency. To fit and validate the framework, we train over 150 VLMs on 34 LLMs spanning 7 model families under a strictly controlled recipe. Evaluations on more than 200 textual and 50 multimodal benchmarks show that the law accurately extrapolates transfer rate from models up to 8B parameters to 72B-scale backbones, predicts full VLM training trajectories with high fidelity, and generalizes to entirely held-out model families. Beyond the scaling law, our analysis surfaces actionable insights: certain textual benchmarks negatively correlate with multimodal performance, exposing latent benchmark-gaming behavior; base LLMs outperform instruction-tuned counterparts as VLM backbones due to higher absorption rates and lower data-scaling decay; and different model families occupy distinct positions in the transfer--absorption space. The framework turns backbone selection from costly empirical sweeps into a principled, quantitative decision. Code and data are available at this https URL.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2608.00013 [cs.CL] |
| (or arXiv:2608.00013v3 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2608.00013 arXiv-issued DOI via DataCite |
Submission history
From: Qiang Wang [view email]
[v1]
Wed, 24 Jun 2026 13:23:33 UTC (540 KB)
[v2]
Mon, 24 Aug 2026 15:47:03 UTC (620 KB)
[v3]
Thu, 8 Oct 2026 09:45:01 UTC (620 KB)
来源:arXiv:cs.CL · arxiv.org