跳到正文
arXiv:cs.AI· Jinghao Xu, Zhenhua Guo, Xiaofeng Zhu, Xiaoshuang Shi·· 6 小时前AI 评分31

ACML:面向多模态学习的动态对齐与校准框架

Dynamic Alignment and Calibration for Multimodal Learning

AI 导读

针对现有多模态学习存在的静态跨模态对齐过度约束、以及置信度感知融合忽略特征量级差异的问题,研究者提出对齐与校准驱动的多模态学习框架 ACML。该框架包含动态跨模态三元组对齐模块,对高置信度正样本对施加强语义一致性约束,并按置信度差距鼓励高低置信度正样本对之间的多样化表示学习;同时引入差异感知注意力校准策略,依据特征量级与置信度差异自适应调整注意力正则化。

正文

View PDF HTML (experimental)

Abstract:Dynamic multimodal learning aims to learn robust representations by adaptively modeling information discrepancies across modalities. However, existing methods still suffer from two limitations: (i) static cross-modal alignment strategies usually impose uniform constraints on all samples while overlooking sample-wise variations, potentially leading to unreasonable over-alignment; and (ii) confidence- or uncertainty-aware fusion methods often fail to adequately account for feature magnitude and confidence differences across modalities. For modality pairs with significant feature magnitude differences or small confidence gaps, it might be unreliable to strictly align fusion weights according to confidence. To address these issues, we propose an Alignment- and Calibration-driven Multimodal Learning framework (ACML). Specifically, ACML incorporates a dynamic cross-modal triplet alignment module, which enforces strong semantic consistency for high-confidence positive pairs while encouraging diverse representation learning between high- and low-confidence positive pairs according to their confidence gaps. Additionally, ACML introduces a difference-aware attention calibration strategy that adaptively adjusts attention regularization based on feature magnitude and confidence differences across modalities, thereby mitigating biases caused by unreasonable fusion constraints. Extensive experiments on multiple multimodal benchmark datasets demonstrate that ACML consistently achieves superior performance and robustness over recent state-of-the-art methods.
Comments: 17 pages
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.07928 [cs.CV]
  (or arXiv:2610.07928v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.07928

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jinghao Xu [view email]
[v1] Tue, 6 Oct 2026 08:05:46 UTC (276 KB)

来源:arXiv:cs.AI · arxiv.org