跳到正文
arXiv:cs.CL· Saptarshi Sengupta, Shuhua Yang, Paul Kwong Yu, Fali Wang, Suhang Wang·· 4 小时前AI 评分34

BioMol-MQA:面向 LLM 生物分子交互推理的多模态问答数据集

BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions

AI 导读

研究者发布 BioMol-MQA,一个针对多重用药(polypharmacy)的多模态问答数据集,由含文本与分子结构的多模态知识图谱(KG)和用于测试 LLM 检索与推理能力的难题两部分组成。基准测试显示,现有 LLM 难以回答这些问题,仅在获得必要背景数据时表现较好,表明需要更强的 RAG 框架。该数据集已被 ICDM 2026 Applied Papers Track 接收。

正文

View PDF HTML (experimental)

Abstract:Retrieval augmented generation (RAG) has shown great power in improving Large Language Models (LLMs). However, most existing RAG-based LLMs are dedicated to retrieving single modality information, mainly text; while for many real-world problems, such as healthcare, information relevant to queries can manifest in various modalities such as knowledge graph, text (clinical notes), and complex molecular structure. Thus, being able to retrieve relevant multi-modality domain-specific information, and reason and synthesize diverse knowledge to generate an accurate response is important. To address the gap, we present BioMol-MQA, a new question-answering (QA) dataset on polypharmacy, which is composed of two parts (i) a multimodal knowledge graph (KG) with text and molecular structure for information retrieval; and (ii) challenging questions that designed to test LLM capabilities in retrieving and reasoning over multimodal KG to answer questions. Our benchmarks indicate that existing LLMs struggle to answer these questions and do well only when given the necessary background data, signaling the necessity for strong RAG frameworks.
Comments: Accepted to ICDM 2026 Applied Papers Track
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2506.05766 [cs.CL]
  (or arXiv:2506.05766v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2506.05766

arXiv-issued DOI via DataCite

Submission history

From: Saptarshi Sengupta [view email]
[v1] Fri, 6 Jun 2025 05:48:22 UTC (2,162 KB)
[v2] Thu, 1 Oct 2026 18:47:08 UTC (487 KB)

来源:arXiv:cs.CL · arxiv.org