arXiv:cs.LG· Mina Abbaszadeh, Matilda Karabina Moore, Raem Haq, Martha Lewis, Mehrnoosh Sadrzadeh·· 4 小时前AI 评分33
Mu-DisCoCat:面向量子处理器组合泛化的变分流水线
Mu-DisCoCat: A Variational Pipeline for Compositional Generalization on Quantum Processors
AI 导读
研究者提出 Mu-DisCoCat 多模态变分量子学习框架,将 DisCoCat 映射到变分量子电路上以实现组合概念泛化(CoCoGen)。该框架先学习单物体图像-文本对的稳定表示,再固定这些表示学习多物体间的关系,在经典模拟下其关系 OOD 准确率超过所评估的 CLIP 基线。
正文
Abstract:Achieving compositional concept generalization (CoCoGen), the ability to understand novel situations by recombining learned primitives, remains a fundamental challenge in artificial intelligence. Compositional semantic models such as Compositional Distributional Semantics (DisCoCat) offer solutions by generalising vectors to tensors, but suffer from scaling bottlenecks when learning the tensors. Mapping DisCoCat onto Variational Quantum Circuits (VQCs) resolves this limitation for text, yet the methodology has not been expanded to multimodal situations such as the ones involved in CoCoGen. This paper introduces Mu-DisCoCat: a multimodal variational quantum learning framework for DisCoCat that achieves CoCoGen. The framework first learns stable object representations from single-object image-text pairs, then fixes these and uses them to learn the relations between them in multi-object situations. In classical simulations, the model used Uhlmann state fidelity to compute the overlap between the multimodal circuit representations and achieved higher relational OOD accuracy than the evaluated CLIP baseline. Its deployment was evaluated using the destructive SWAP test across noisy quantum emulators, including a range of IBM fake backends, IQM FakeAphrodite, and the IBM Marrakesh quantum processor. Despite real-world device noise, the hardware-executed models maintained a strong positive correlation with simulated fidelities, reliably distinguishing unseen similar and dissimilar pairs. Our work establishes a framework for executing CoCoGen on VQCs, demonstrating a viable use case for near-term quantum hardware.
| Subjects: | Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2610.08131 [cs.LG] |
| (or arXiv:2610.08131v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08131 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mina Abbaszadeh [view email]
[v1]
Tue, 6 Oct 2026 10:45:43 UTC (2,244 KB)
来源:arXiv:cs.LG · arxiv.org