arXiv:cs.AI· Lingli Ge, Yubin Wang, Junyuan Gao, Jiahe Song, Jiaxing Sun, Boyu Zhu, Haote Yang, Jingchao Wang, Lixin Ma, Jiang Wu, Yuqiang Li, Conghui He·· 4 小时前AI 评分48
RxnOptBench:面向有机方法学反应条件优化的 LLM 基准测试
RxnOptBench: Benchmarking LLMs for Reaction-Condition Optimization in Organic Methodology
AI 导读
研究者推出 RxnOptBench,用于评测 LLM 从真实条件筛选表中选出最优反应条件的能力,题目取自 2025 年有机方法学论文的优化表格,并按结合收率与 ee、dr、rr 的连续相对分数评分。
正文
Authors:Lingli Ge, Yubin Wang, Junyuan Gao, Jiahe Song, Jiaxing Sun, Boyu Zhu, Haote Yang, Jingchao Wang, Lixin Ma, Jiang Wu, Yuqiang Li, Conghui He
Abstract:Chemical reaction-condition optimization -- choosing the catalyst, ligand, solvent, reagent, temperature, time, and atmosphere that jointly maximize yield and stereoselectivity -- is a central, judgement-laden subtask of organic methodology research that large language models are increasingly expected to support. Yet existing chemistry benchmarks evaluate reaction-class labelling, retrosynthesis, or SMILES manipulation, and do not ask models to read a real condition-screening table and pick the best set. We introduce RxnOptBench, a benchmark whose every option and precedent is a real wet-lab entry mined from the optimization tables of organic-methodology papers published in 2025, graded by a continuous relative score derived from a declared headline utility that combines reported yield with enantiomeric excess (ee), diastereomeric ratio (dr), and regioisomeric ratio (rr), and equipped with a paired precedents-vs-no-precedents design that isolates in-context use of literature evidence from parametric memorization. Across nine frontier LLMs and three Chemistry LLMs, even the best models leave substantial headroom: chemistry-specialized models fall to the random-baseline floor on multi-axis selection, while open-weight models have closed most of the gap to proprietary frontier models. We release the final human-reviewed benchmark test set and evaluation code.
| Comments: | Accepted to NeurIPS 2026 (Evaluations & Datasets Track) |
| Subjects: | Chemical Physics (physics.chem-ph); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.02242 [physics.chem-ph] |
| (or arXiv:2610.02242v1 [physics.chem-ph] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02242 arXiv-issued DOI via DataCite |
Submission history
From: Lingli Ge [view email]
[v1]
Tue, 29 Sep 2026 15:17:31 UTC (6,571 KB)
来源:arXiv:cs.AI · arxiv.org