arXiv:cs.AI· Youngtaek Oh, Qiyu Wu, Hiromi Wakaki, Junmo Kim, Yuki Mitsufuji·· 3 小时前
Syn-Omni:面向全模态嵌入的结构化专精与渐进协作框架
Syn-Omni: Structured Specialization and Progressive Collaboration for Omnimodal Embeddings
AI 导读
Syn-Omni 提出 OME-LoRA 与 PSR 机制,将全模态嵌入适配拆分为共享 LoRA 路径与模态专家 LoRA 路径,并让专家先建立模态先验再进行跨模态协作。该框架在覆盖图像、视频、音频和视听模态的 81 项任务上评测,持续优于全模态基线,论文已被 EMNLP 2026 接收。
正文
Abstract:Omnimodal embeddings naturally involve both shared representations and modality-specific features across heterogeneous inputs. However, existing omnimodal embedding methods often rely on a single shared parameter space over mixed-modality data, limiting structural separation between universal and modality-specific representations. To address this, we propose Syn-Omni, a unified framework for structured omnimodal adaptation with modality specialization and controlled cross-modal collaboration. Specifically, we introduce Orthogonal Modality-Expert LoRA (OME-LoRA), which decomposes adaptation into a shared LoRA path for universal semantics and modality-expert LoRA paths for modality-aware specialization. Furthermore, Progressive Synergy Routing (PSR) enables experts to first establish modality-specific priors, then gradually interact with other modality-experts for cross-modal synergy. Evaluated across 81 diverse tasks spanning image, video, audio, and audiovisual modalities, Syn-Omni consistently outperforms omnimodal baselines, demonstrating the effectiveness of structured specialization and cross-modal progressive collaboration.
| Comments: | Accepted to EMNLP 2026 (Long, Findings). Code: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR) |
| Cite as: | arXiv:2610.12256 [cs.CV] |
| (or arXiv:2610.12256v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12256 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Youngtaek Oh [view email]
[v1]
Thu, 8 Oct 2026 16:26:03 UTC (666 KB)
来源:arXiv:cs.AI · arxiv.org