arXiv:cs.LG· Yang Qiao, Yuntong Hu, Bowen Zhu, Hasibul Haque, Liang Zhao·· 5 小时前AI 评分30
RCML:以语义关系为条件的多模态表示学习框架
Multimodal Representation Learning Conditioned on Semantic Relations
AI 导读
研究者提出关系条件化多模态学习框架 RCML,将自然语言描述语义关系作为显式条件,使同一样本在不同关系语境下获得不同表示,而非像 CLIP 那样只输出单一关系无关的嵌入向量。该框架构建关系感知训练对,引入关系条件化模块,并用统一对比目标联合建模跨模态对齐与关系诱导的样本间结构。在多个数据集上,RCML 在零样本、微调和域外设置的检索与分类任务中均持续优于强基线。
正文
Abstract:Multimodal representation learning has been largely driven by contrastive models such as CLIP, which learn a shared embedding space by aligning paired image-text samples. While effective for general-purpose representation learning, such models typically produce a single embedding per sample that is reused across different semantic relations and contexts. However, in many real-world applications, relevance between samples is inherently relation-dependent, with different semantic relations emphasizing different aspects of multimodal data.
In this work, we propose Relation-Conditioned Multimodal Learning (RCML), a framework that treats semantic relations as explicit conditions of multimodal representation learning. Rather than producing relation-agnostic embeddings, RCML learns representations conditioned on natural-language relation descriptions, allowing the same sample to be represented differently under different relational contexts. The framework constructs relation-aware training pairs, introduces a relation-conditioned module to adapt embeddings to relation semantics, and employs a unified contrastive objective to jointly model cross-modal alignment and relation-induced inter-sample structure.
Experiments on multiple datasets show that RCML consistently outperforms strong baselines on retrieval and classification tasks in zero-shot, fine-tuned, and out-of-domain settings, highlighting the effectiveness of leveraging semantic relations to guide multimodal representation learning.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2508.17497 [cs.LG] |
| (or arXiv:2508.17497v3 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2508.17497 arXiv-issued DOI via DataCite |
Submission history
From: Yang Qiao [view email]
[v1]
Sun, 24 Aug 2025 19:36:18 UTC (4,623 KB)
[v2]
Sat, 9 May 2026 13:53:28 UTC (4,834 KB)
[v3]
Thu, 1 Oct 2026 19:34:38 UTC (2,844 KB)
来源:arXiv:cs.LG · arxiv.org