MBZUAI 发布 Omni-Embed-Mini:0.9B 参数多模态嵌入模型不遗忘文本检索
Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
MBZUAI 发布 Omni-Embed-Mini,0.9B 参数模型将文本、语音、音频、图像、视频和富视觉文档映射到同一共享余弦空间,且文本侧参数完全不动。
Published on Oct 1
Authors:
,
,
,
Abstract
Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space without updating any text-side parameter. Our key insight is that the teacher signal requires no separate embedding model: each media sample is paired with a dense cascaded caption, and the teacher target is simply the frozen backbone's own embedding of that caption. Because teacher and student share the same backbone weights, they inhabit byte-identical geometry, and lightweight projectors plus phased LoRA adapters on the modality encoders suffice for alignment. Training combines a Matryoshka SigLIP contrastive loss with an online hybrid hard-negative miner whose negatives sharpen as the encoder improves. The recipe carries over to a 2.3B variant by swapping in a native vision-language backbone. Omni-Embed-Mini-0.9B keeps its text weights bit-identical to the backbone, so training cannot regress text retrieval (49.57 nDCG@10 on MTEB-v2 BEIR-8), while extending it to five additional modalities, and is ~2.7x to 9.5x smaller than every open omni embedder we compare against. The 2.3B variant is competitive with the closed gemini-embedding-2, edging ahead of it on the overall-modality average. Models, code, data and evaluation harness are on our project page: https://omniembed.cvmbzuai.com
View arXiv page View PDF Project page GitHub Add to collection
Get this paper in your agent:
hf papers read 2610.02148
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this paper 3
Feature Extraction • Updated about 1 hour ago • 28 • 3MBZUAI/Omni-Embed-Mini-0.9B
Feature Extraction • Updated about 1 hour ago • 26 • 1MBZUAI/Omni-Embed-Mini-2.3B
Updated about 4 hours ago • 24MBZUAI/Omni-Embed-Mini-0.9B-onnx
Datasets citing this paper 1
MBZUAI/Omni-Sets
Viewer • Updated about 1 hour ago • 591k • 1.6k • 1
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2610.02148 in a Space README.md to link it from this page.
Collections including this paper 1
来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co