arXiv:cs.AI(全量分类)· Chenmu Zhang, Levi Felix, Jun-Jie Zhang, Xingfu Li, Xuelian Jiang, Tao Jiang, Subhendu Mishra, Xixi Qin, Boris Yakobson·· 5 小时前AI 评分44
CompMat-Bench:面向计算材料科学的 AI 智能体基准测试
CompMat-Bench: Benchmarking AI Agents for Computational Materials Science
AI 导读
研究者推出 CompMat-Bench,一个源自近期计算材料研究、共 94 项任务的基准,通过预先复现研究步骤、让智能体准备输入并分析输出,避免评测时运行昂贵模拟。基准支持单任务与工作流、完整与精简方法指导四种条件,三个 LLM 智能体在单任务完整指导下通过率为 66.0-90.4%。更长工作流与精简指导会以不同方式限制智能体表现,失败多归因于科学错误而非软件使用错误。
正文
Abstract:Evaluating AI agents on scientific research tasks is constrained by the time and resources required for the underlying experiments or calculations. In computational materials research, repeating the same expensive simulations across agents and trials can make evaluation impractical. We introduce CompMat-Bench, a benchmark of 94 tasks derived from recently published computational materials studies, each asking agents to complete a step toward achieving the study's scientific goal. We reproduce the research steps in advance and assess agents on preparing inputs and analyzing outputs for expensive simulations, so expensive simulations can be avoided during evaluation. The reproduced inputs and results serve as ground truth for grading agents with fixed rules, without an LLM judge. The benchmark supports four evaluation conditions: single tasks and workflows composed of related tasks, each with full or reduced methodological guidance. With full guidance on single tasks, agents based on three LLMs demonstrate the ability to complete individual materials research steps, with pass rates of 66.0-90.4% across 94 tasks. Both longer workflows and reduced guidance can limit agent performance, but in different ways for different agents: they lower the pass rates of the weaker agents, whereas the strongest agent falls only when a long workflow is combined with reduced guidance. Failure analysis attributes most failures to scientific errors rather than to errors in software usage. CompMat-Bench provides a basis for comparing agents on the steps of real materials research and for analyzing agent failure modes.
| Subjects: | Artificial Intelligence (cs.AI); Materials Science (cond-mat.mtrl-sci); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00636 [cs.AI] |
| (or arXiv:2610.00636v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00636 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Chenmu Zhang [view email]
[v1]
Wed, 30 Sep 2026 19:37:09 UTC (275 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org