arXiv:cs.AI(全量分类)· Nishtha N. Vaidya, Stephan Grimm, Thomas Hubauer, Thomas A. Runkler·· 5 小时前AI 评分39
基于本体、经推理器验证的 LLM 科学推理评测基准
Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI
AI 导读
研究者提出一条自动生成流水线,可从任意公理化的 OWL 2 本体构建多选题基准,正确答案由本体保证,干扰项通过扰动类定义公理右侧类表达式生成并经 OWL 推理器蕴含检查验证其错误性。在 Pizza、PMDco、DOID 三个本体上分别生成 112、2,491、15,216 道题,六款 LLM 零样本准确率为 41.1%-76.8%,高于 25% 随机基线。
正文
Abstract:Large language models (LLMs) increasingly underpin scientific AI applications that reason over structured knowledge, from biomedical question answering to materials informatics. However, their logical reasoning often falls short, producing factual inaccuracies unacceptable in these settings. Reliable evaluation remains challenging: manual dataset construction scales poorly, and LLM-based generation risks embedding the very flaws it aims to measure. High-quality benchmarks must ground both correct and incorrect labelled examples in explicit background knowledge, formally verifiable by a standard reasoner. We propose a pipeline that automatically generates ontology-grounded multiple-choice question (MCQ) benchmarks from any sufficiently axiomatised OWL 2 ontology, with correct answers grounded in the ontology by design. Distractors are generated by perturbing the right-hand-side class expressions of class definition axioms, and their incorrectness is formally verified by an OWL reasoner via entailment checks. We evaluate the pipeline on three ontologies: Pizza (small, academic), PMDco (complex, materials science), and DOID (large, biomedical), generating 112, 2,491, and 15,216 MCQs respectively. Distractors span four semantic categories from class unsatisfiability to weakened subsumptions, enabling diagnostic evaluation of specific reasoning failures. Items meet natural language quality standards: mean LLM judge scores of 4.02, 4.36, and 3.36 out of 5 confirm fluency, and correct-answer-to-distractor similarity above 0.8 shows that wrong options cannot be dismissed on surface form alone. Six LLMs evaluated zero-shot achieve 41.1-76.8% accuracy, well above the 25% random-guessing baseline, indicating the benchmarks are challenging and discriminative. This work is a step towards more reliable benchmarks for assessing logical reasoning in scientific AI.
| Comments: | 17 pages, 2 figures. Accepted at the AI Data Readiness for Scientific Discovery (AIDaR) Workshop at the 40th Conference on Neural Information Processing Systems (NeurIPS 2026), Paris |
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00682 [cs.AI] |
| (or arXiv:2610.00682v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00682 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Nishtha Namdeo Vaidya [view email]
[v1]
Wed, 30 Sep 2026 20:23:55 UTC (420 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org