arXiv:cs.AI· Xuanming Cui, Shlok Kumar Mishra, Wentao Bao, Aashu Singh, Zihao Wang, Xiangjun Fan, Jun Xiao, Ser-Nam Lim, Jianpeng Cheng·· 3 小时前
MoEMB:用高效 MoE 模型扩展通用多模态嵌入
MoEMB: Scaling Universal Multimodal Embeddings with Efficient Mixture-of-Experts Models
AI 导读
MoEMB 提出沿专家轴扩展通用多模态嵌入(UME),通过混合专家(MoE)在保留单向量、非自回归编码的同时提升编码器容量。在公开 MMEB 系列数据训练下,MoEMB 以仅 3B 激活参数在 MMEB-V2 与 MRMR 上刷新 SOTA,超越激活参数超 4 倍的 TTE 方法且计算量显著更少。该工作还首次系统研究了 MoE 嵌入的自适应计算,涵盖基于训练与仅推理的多类策略。
正文
Abstract:Universal multimodal embedding (UME) increasingly demands encoder's capacity for handling a broad range of tasks and modalities with increased complexity. Prior scaling methods either increase the representation size, retrieval effort, or scales the encoder into a heavy multimodal LLM. Recent works, such as Think-Then-Embed (TTE), explore scaling via reasoning tokens. However, embedding models are hard to scale up: increasing parameters directly tradeoffs for the large training batch size that contrastive learning needs, and retrieval has to be served under tight latency. Moreover, UME tasks are diverse in complexity, where scaling up embedders can bring significant redundant computation. In this work, we propose MOEMB, which instead scales UME along the expert axis through mixture-of-experts (MoE), growing encoder capacity while preserving single-vector, non-autoregressive encoding. Through a systematic study of the design space and training recipes for MoE-based UME, MoEMB sets a new state of the art on both MMEB-V2 and MRMR among models trained on public MMEB-family data: with only 3B active parameters, MoEMB surpasses TTE-based methods with >4x active parameters, using significantly less computes. To further improve the scalability and efficiency, we conduct the first comprehensive study of adaptive computation for MoE-based embedding, spanning diverse strategies across training-based and inference-only methods. Together, these results support expert scaling as an effective and efficient direction for UME, with adaptive computation further improving efficiency for MLLM-based embedding models towards large-scale retrieval and recommendation systems.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.08663 [cs.LG] |
| (or arXiv:2609.08663v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.08663 arXiv-issued DOI via DataCite |
Submission history
From: Xuanming Cui [view email]
[v1]
Tue, 8 Sep 2026 12:29:50 UTC (3,431 KB)
[v2]
Wed, 7 Oct 2026 18:58:19 UTC (3,443 KB)
来源:arXiv:cs.AI · arxiv.org