arXiv:cs.AI· Donghyun Lee, Taehoon Lee, Geonhee Ahn, Jieun Kim, Jihyun Park, Suyeon Cho, Yoona Kim, Chaerim Shin, Hoi Ri Moon, Jonggeol Na, Sukho Hong, Jihwan Oh, Soo Kyung Kim·· 5 小时前AI 评分34
MOF-Verify:面向 MOF 假设验证的失败感知智能体框架
MOF-VERIFY: A Failure-Aware Agentic Harness for MOF Hypothesis Verification
AI 导读
研究者提出 MOF-Verify,一个失败感知的智能体框架,针对结构、文献、证据充分性和计算四类瓶颈进行验证后再给出最终结论。配套诊断基准含四类任务,其中 T-MOF-1-3 在闭卷、检索增强和 oracle 证据三种设置下定位知识获取、证据获取与推理失败,T-MOF-4 单独评估基于 MLIP 的计算验证。
正文
Authors:Donghyun Lee, Taehoon Lee, Geonhee Ahn, Jieun Kim, Jihyun Park, Suyeon Cho, Yoona Kim, Chaerim Shin, Hoi Ri Moon, Jonggeol Na, Sukho Hong, Jihwan Oh, Soo Kyung Kim
Abstract:Large language models are increasingly used as reasoning components in AI-driven materials Co-Scientists, yet the reliability of the resulting verification pipeline remains unclear. Metal-organic frameworks (MOFs) provide a particularly challenging setting because structures may appear under different identifiers, synthesis outcomes depend strongly on experimental conditions, evidence is distributed across heterogeneous sources, and some hypotheses require computation rather than literature alone. We introduce a diagnostic benchmark with four task families covering structural grounding, synthesis-condition verification, evidence-sufficiency verification, and MLIP-based computational verification. T-MOF-1-3 are evaluated under closed-book, retrieval-enabled, and oracle-evidence settings to localize failures in knowledge access, evidence acquisition, and reasoning, while T-MOF-4 separately evaluates computational verification. Guided by these diagnosed failure modes, we develop MOF-Verify, a failure-aware agentic harness that targets structural, literature, evidence-sufficiency, and computational bottlenecks before producing a final verdict. Across multiple backbone LLMs, MOF-Verify substantially improves hypothesis-verification performance over direct inference and retrieval-based baselines. Benchmark datasets are released at this https URL.
| Comments: | Accepted at the NeurIPS 2026 Workshops XAI4Science and AI4Mat |
| Subjects: | Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE) |
| Cite as: | arXiv:2610.03056 [cs.AI] |
| (or arXiv:2610.03056v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03056 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Dong Hyun Lee [view email]
[v1]
Fri, 2 Oct 2026 09:36:55 UTC (324 KB)
来源:arXiv:cs.AI · arxiv.org