跳到正文
原文
HuggingFace Daily Papers(社区热门论文)·· 11 小时前AI 评分55

MBZUAI 发布 Omni-Embed-Mini:0.9B 参数多模态嵌入模型不遗忘文本检索

Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation

AI 导读

MBZUAI 发布 Omni-Embed-Mini,0.9B 参数模型将文本、语音、音频、图像、视频和富视觉文档映射到同一共享余弦空间,且文本侧参数完全不动。

正文

Published on Oct 1

Authors:

,

,

,

Abstract

Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space without updating any text-side parameter. Our key insight is that the teacher signal requires no separate embedding model: each media sample is paired with a dense cascaded caption, and the teacher target is simply the frozen backbone's own embedding of that caption. Because teacher and student share the same backbone weights, they inhabit byte-identical geometry, and lightweight projectors plus phased LoRA adapters on the modality encoders suffice for alignment. Training combines a Matryoshka SigLIP contrastive loss with an online hybrid hard-negative miner whose negatives sharpen as the encoder improves. The recipe carries over to a 2.3B variant by swapping in a native vision-language backbone. Omni-Embed-Mini-0.9B keeps its text weights bit-identical to the backbone, so training cannot regress text retrieval (49.57 nDCG@10 on MTEB-v2 BEIR-8), while extending it to five additional modalities, and is ~2.7x to 9.5x smaller than every open omni embedder we compare against. The 2.3B variant is competitive with the closed gemini-embedding-2, edging ahead of it on the overall-modality average. Models, code, data and evaluation harness are on our project page: https://omniembed.cvmbzuai.com

View arXiv page View PDF Project page GitHub Add to collection

Get this paper in your agent:

hf papers read 2610.02148

Don't have the latest CLI?

curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 3

MBZUAI/Omni-Embed-Mini-0.9B

Feature Extraction •

Updated about 1 hour ago

•

28

•

3

MBZUAI/Omni-Embed-Mini-2.3B

Feature Extraction •

Updated about 1 hour ago

•

26

•

1

MBZUAI/Omni-Embed-Mini-0.9B-onnx

Updated about 4 hours ago

•

24

Datasets citing this paper 1

MBZUAI/Omni-Sets

Viewer •

Updated about 1 hour ago

•

591k

•

1.6k

•

1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.02148 in a Space README.md to link it from this page.

Collections including this paper 1

来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co