arXiv:cs.AI· Qun Dai, Liangjian Wen, Jiang Duan, Yong Dai, Dongkai Wang, Maolin Wang, Mingjie Wang, Jianzhuang Liu, He Yan, Zhao Kang·· 3 小时前
HRIL:用高阶张量建模学习多模态协同信息
HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling
AI 导读
针对多模态协同信息难以捕捉的问题,研究者提出 HRIL,通过对模态嵌入构建经验交叉矩张量并用 Tucker 分解提取核心张量,配合协同感知正则项保留高阶耦合能力。在受控协同任务和真实基准上,HRIL 相比现有多模态对比方法取得一致提升,协同交互主导的任务上增益尤为明显。代码已开源,论文被 NeurIPS 2026 接收。
正文
Abstract:Self-supervised multimodal representation learning has achieved remarkable success across diverse domains, yet capturing synergistic information remains challenging due to the complexity of cross-modal interactions. Unlike the shared information across individual modalities, synergy arises when task-relevant signals emerge only from the joint configuration of multiple modalities and cannot be recovered from any modality in isolation. This work focuses on how to preserve the information capacity for such synergistic signals in multimodal representations. The key observation is that synergistic information is reflected in higher-order statistical dependence among modalities, which provides a principled target for explicitly modeling joint interactions. Motivated by this insight, we propose Higher-order Representation and Information Learning (HRIL), which constructs an empirical cross-moment tensor over modality embeddings to represent multi-way interactions. HRIL employs Tucker decomposition to obtain a core tensor, complemented by a synergy-aware regularizer that prevents energy concentration and preserves higher-order coupling capacity for synergistic information capture. Experiments on the controlled synergy task and real-world benchmarks demonstrate consistent improvements over existing multimodal contrastive methods, with notable gains on tasks dominated by synergistic interactions. Code is released at this https URL.
| Comments: | Accepted at NeurIPS 2026 |
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.12393 [cs.AI] |
| (or arXiv:2610.12393v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12393 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Qun Dai [view email]
[v1]
Thu, 8 Oct 2026 17:39:45 UTC (260 KB)
来源:arXiv:cs.AI · arxiv.org