跳到正文
原文
arXiv:cs.AI(全量分类)· Xin Wang, Zichuan Ying, Xinna Lin, Junqi Zhang, Hanyi Xiong, Tianyu Gao, Hairong Zhang, Qixiang Hua, Botian Shi, Zhenhailong Wang, Kaicheng Yu·· 5 小时前AI 评分39

R-GroundBench:面向 Markush 分子编辑中 R 基团定位的诊断基准

R-GroundBench: A Diagnostic Benchmark for R-Group Groundingin Markush Molecular Editing

AI 导读

研究者提出 R-GroundBench,一个基于真实专利 Markush 结构构建的诊断基准,包含难度与模态可控的多选题(VQA)赛道和开放式生成赛道,用于评估 R 基团定位能力。

正文

Authors:Xin Wang, Zichuan Ying, Xinna Lin, Junqi Zhang, Hanyi Xiong, Tianyu Gao, Hairong Zhang, Qixiang Hua, Botian Shi, Zhenhailong Wang, Kaicheng Yu

View PDF HTML (experimental)

Abstract:Recent advances in AI for scientific discovery enable molecular understandingand design, yet reasoning over incomplete chemical representations this http URL structures, which encode molecular families through variable R-groupplaceholders (\textit{R\textsubscript{1}}, \textit{R\textsubscript{2}}, \textit{X}, etc.), are ubiquitous in pharmaceutical patents and requiregrounding across molecular, textual, and chemical this http URL, existing molecule-language benchmarks focus on fully specifiedmolecules, leaving R-group grounding largely this http URL introduce R-GroundBench:, a diagnostic benchmark built from real patent Markushstructures, featuring a Multiple-Choice (VQA) track with controlled difficultyand modality splits, and an open-ended Generation this http URL results reveal a substantial gap between recognition andmolecular this http URL models achieve over 90\% accuracy on Easy VQA, performance drops to56--66\% on Hard VQA when shortcuts are this http URL-domain VLMs also remain unreliable, achieving only 25.7--46.2\% on HardVQA despite domain-specific this http URL, Generation Exact Match remains below 20\% for most models and below8\% when visual input is this http URL findings reveal that current AI systems lack reliable grounding andexecution for Markush editing, highlighting challenges for AI-drivenscientific discovery.
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.00700 [cs.AI]
  (or arXiv:2610.00700v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.00700

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xin Wang [view email]
[v1] Wed, 30 Sep 2026 20:45:31 UTC (25,519 KB)

来源:arXiv:cs.AI(全量分类) · arxiv.org