跳到正文
arXiv:cs.AI· Srikar Alla, Ali Shiri Sichani, Chi-Ren Shyu·· 6 小时前AI 评分34

QEMFN:基于可训练纠缠的资源感知混合视觉语言融合网络

Quantum Entangled Multimodal Fusion Networks (QEMFN): Resource-Aware Hybrid Vision-Language Fusion via Trainable Entanglement

AI 导读

研究者提出量子纠缠多模态融合网络(QEMFN),一种以参数化纠缠作为结构化归纳偏置的混合量子-经典框架,用于视觉语言融合。在参数预算匹配且冻结相同 CLIP 主干下,QEMFN 在 COCO-5k 和 Flickr30k 上优于 MLP、张量融合、FiLM、交叉注意力等经典基线。该框架在基于 shot 的估计、噪声建模模拟后端及真实超导设备上运行,论文明确不主张量子计算优势。

正文

View PDF HTML (experimental)

Abstract:Multimodal vision-language systems typically fuse image and text embeddings through classical operators such as concatenation, attention, bilinear pooling, or tensor interactions. We propose Quantum Entangled Multimodal Fusion Networks (QEMFN), a hybrid quantum-classical framework that introduces parameterized entanglement as a structured inductive bias for multimodal fusion. Pretrained visual and textual features are projected into compact latent spaces, encoded as angle-parameterized quantum states, processed through intra-modal and paired cross-modal entangling circuits, and measured to produce fused representations for retrieval. Under matched parameter budgets and identical frozen CLIP backbones, QEMFN outperforms classical fusion baselines on COCO-5k and Flickr30k, including multilayer perceptron, tensor fusion, FiLM, cross-attention, compact transformer, and a dequantized paired-topology analogue. An ablation suite isolates the quantum module's contribution from the surrounding classical projections, and quantum-centric analyses report Meyer-Wallach entangling capability, expressibility, gradient variance against barren-plateau bounds, and entropy-performance correlation under controls for training progress alongside an intervention study on the entangling component. QEMFN is executed under shot-based estimation, a noise-modeled fake backend, and a real superconducting device with zero-noise extrapolation. This work does not claim quantum computational advantage; the contribution is the framework together with a controlled empirical and quantum-centric evaluation that positions trainable entanglement as an interpretable, hardware-executable fusion mechanism at scales accessible on contemporary devices.
Subjects: Artificial Intelligence (cs.AI); Quantum Physics (quant-ph)
Cite as: arXiv:2610.08216 [cs.AI]
  (or arXiv:2610.08216v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.08216

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ali Shiri Sichani [view email]
[v1] Tue, 6 Oct 2026 12:08:58 UTC (408 KB)

来源:arXiv:cs.AI · arxiv.org