跳到正文
arXiv:cs.LG· Peilin Yang, Xiaoyu Liu, Jian Sun, Qinghua Tao·· 3 小时前AI 评分34

通过核典型相关分析重新审视 VLM 的视觉表征增强

Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis

AI 导读

研究提出基于核典型相关分析(KCCA)的特征子空间对齐方法,并扩展为三视图 3vKCCA,将预训练文本编码器的投影纳入统一优化框架。在 CLIP ViT-L/14 与 ImageNet-1K 上,3vKCCA 将 MMVP-VLM 准确率从 17.8 提升至 25.9,同时保持零样本图文检索性能。该方法基于 KKT 条件推导出高效端到端训练方案,避免求解 KCCA 的特征值问题。

正文

View PDF HTML (experimental)

Abstract:Vision-language models such as CLIP exhibit strong semantic generalization, but remain limited in fine-grained visual perception. A recent work named KUEA presents a natural remedy by finetuning the image encoder under the supervision of the vision-centric DINOv2 to align their kernel matrices element-wisely, while regularizing the embeddings to remain close to the pretrained visual encoder for preserving image-text semantics in CLIP. However, we show that diminishing the role of the alignment loss to DINOv2 does not necessarily degrade its fine-grained visual performance, suggesting that the kernel-matrix discrepancy may be insufficient for further visual representation enhancement, motivating us to revisit the alignment formulation. In this work, we present a novel perspective to characterize representation alignment on feature subspaces through Kernel Canonical Correlation Analysis (KCCA), which maximizes the projection correlations. In optimization, we derive an efficient end-to-end training scheme upon KKT conditions, avoiding the eigenvalue problem in KCCA. Further, we extend our method into a 3-view formulation, i.e., 3vKCCA, in which the projections from the pretrained text encoder are also incorporated under a unified optimization framework for joint alignment. With CLIP ViT-L/14 on ImageNet-1K, our 3vKCCA improves the MMVP-VLM accuracy from 17.8 to 25.9, substantially outperforming the existing methods, and meanwhile maintains zero-shot image--text retrieval performance.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.02718 [cs.CV]
  (or arXiv:2610.02718v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.02718

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Peilin Yang [view email]
[v1] Fri, 2 Oct 2026 02:54:59 UTC (460 KB)

来源:arXiv:cs.LG · arxiv.org