arXiv:cs.LG· Hebao Zhu, Dongxia Wu·· 3 小时前AI 评分34
MLLMs 中的架构依赖融合路径:拼接架构与原生多模态架构对比研究
Architecture-Dependent Fusion Pathways in MLLMs
AI 导读
研究对比了拼接架构与原生多模态两类 MLLMs,通过对齐解耦、注意力路由与熵、内在维度三项递进分析及因果干预实验,揭示出两条不同的融合路径:拼接模型遵循"文本优先、视觉在后"的路径,原生模型则表现出更早的视觉-文本协同适应与特征空间重组。研究还以 visual CKA 作为补充分析检验了柏拉图表示假说,为多模态融合提供了机制性视角,并支持架构感知的多模态表示诊断。
正文
Abstract:Multimodal Large Language Models (MLLMs) achieve strong performance across vision-language tasks, yet the internal mechanisms by which visual and textual information are fused across layers remain insufficiently understood. We investigate representative MLLMs from two architectural paradigms: concatenation architectures and native multimodal architectures. We conduct three progressively connected analyses: alignment decoupling identifies which modality changes, attention routing and entropy characterize how cross-modal information is distributed, and intrinsic dimensionality examines how fusion reshapes feature spaces. Separately, we perform causal intervention experiments as a validation of the resulting interpretation. As a supplementary analysis, we use visual CKA to examine the Platonic Representation Hypothesis. Together, these analyses reveal two distinct fusion pathways: concatenation models follow a text-first, vision-later pathway, whereas native models exhibit earlier visual-textual co-adaptation and feature-space reorganization. This work provides a mechanistic perspective for understanding multimodal fusion and supports architecture-aware diagnostics of multimodal representations.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.03289 [cs.LG] |
| (or arXiv:2610.03289v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03289 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Dongxia Wu [view email]
[v1]
Fri, 2 Oct 2026 13:32:45 UTC (1,915 KB)
来源:arXiv:cs.LG · arxiv.org