跳到正文
arXiv:cs.LG· Hebao Zhu, Dongxia Wu·· 3 小时前AI 评分34

MLLMs 中的架构依赖融合路径:拼接架构与原生多模态架构对比研究

Architecture-Dependent Fusion Pathways in MLLMs

AI 导读

研究对比了拼接架构与原生多模态两类 MLLMs,通过对齐解耦、注意力路由与熵、内在维度三项递进分析及因果干预实验,揭示出两条不同的融合路径:拼接模型遵循"文本优先、视觉在后"的路径,原生模型则表现出更早的视觉-文本协同适应与特征空间重组。研究还以 visual CKA 作为补充分析检验了柏拉图表示假说,为多模态融合提供了机制性视角,并支持架构感知的多模态表示诊断。

正文

View PDF HTML (experimental)

Abstract:Multimodal Large Language Models (MLLMs) achieve strong performance across vision-language tasks, yet the internal mechanisms by which visual and textual information are fused across layers remain insufficiently understood. We investigate representative MLLMs from two architectural paradigms: concatenation architectures and native multimodal architectures. We conduct three progressively connected analyses: alignment decoupling identifies which modality changes, attention routing and entropy characterize how cross-modal information is distributed, and intrinsic dimensionality examines how fusion reshapes feature spaces. Separately, we perform causal intervention experiments as a validation of the resulting interpretation. As a supplementary analysis, we use visual CKA to examine the Platonic Representation Hypothesis. Together, these analyses reveal two distinct fusion pathways: concatenation models follow a text-first, vision-later pathway, whereas native models exhibit earlier visual-textual co-adaptation and feature-space reorganization. This work provides a mechanistic perspective for understanding multimodal fusion and supports architecture-aware diagnostics of multimodal representations.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.03289 [cs.LG]
  (or arXiv:2610.03289v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.03289

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Dongxia Wu [view email]
[v1] Fri, 2 Oct 2026 13:32:45 UTC (1,915 KB)

来源:arXiv:cs.LG · arxiv.org