跳到正文
arXiv:cs.CL· Ziran Li, Qiang Wang, Zhengyu Chen, Shanglin Lei, Borun Chen, Jingang Wang, Xunliang Cai·· 4 小时前

从文本到视觉能迁移什么?VLM 的能力缩放定律与迁移动态

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs

AI 导读

研究者提出 Capability-Driven Multimodal Scaling Law,首个跨模型家族框架,可从 LLM 文本基准的 PCA 能力分数 S 预测 VLM 基准准确率。

正文

View PDF HTML (experimental)

Abstract:Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains fundamentally unprincipled: compute-based scaling laws fail to generalize across model families, and no framework exists for directly predicting VLM performance before training begins. We propose the Capability-Driven Multimodal Scaling Law, the first cross-family framework that predicts VLM benchmark accuracy from directly observable textual capability. Given a low-dimensional capability score $S$ extracted from LLM textual benchmarks via PCA, we model VLM performance as a function of $S$, with a per-backbone transfer rate and an absorption rate that quantifies data-scaling efficiency. To fit and validate the framework, we train over 150 VLMs on 34 LLMs spanning 7 model families under a strictly controlled recipe. Evaluations on more than 200 textual and 50 multimodal benchmarks show that the law accurately extrapolates transfer rate from models up to 8B parameters to 72B-scale backbones, predicts full VLM training trajectories with high fidelity, and generalizes to entirely held-out model families. Beyond the scaling law, our analysis surfaces actionable insights: certain textual benchmarks negatively correlate with multimodal performance, exposing latent benchmark-gaming behavior; base LLMs outperform instruction-tuned counterparts as VLM backbones due to higher absorption rates and lower data-scaling decay; and different model families occupy distinct positions in the transfer--absorption space. The framework turns backbone selection from costly empirical sweeps into a principled, quantitative decision. Code and data are available at this https URL.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2608.00013 [cs.CL]
  (or arXiv:2608.00013v3 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2608.00013

arXiv-issued DOI via DataCite

Submission history

From: Qiang Wang [view email]
[v1] Wed, 24 Jun 2026 13:23:33 UTC (540 KB)
[v2] Mon, 24 Aug 2026 15:47:03 UTC (620 KB)
[v3] Thu, 8 Oct 2026 09:45:01 UTC (620 KB)

来源:arXiv:cs.CL · arxiv.org