arXiv:cs.AI· David Li, Shaamil Karim, Christian Gensbigler·· 6 小时前AI 评分43
BioStudyBench:评估 AI 智能体复现知识截止后生物医学研究的能力
BioStudyBench: Evaluating Agents on Post-Cutoff Biomedical Studies
AI 导读
研究团队推出 BioStudyBench,包含 25 项长周期分析任务,取自 2026 年 7 至 9 月发表、晚于被测模型知识截止时间的生物医学研究,由 404,019 条 PubMed 记录半自动筛选而来。
正文
Abstract:We evaluate whether AI agents can match the reported findings of published biomedical studies using public data. Existing evaluations do not consistently separate analysis from prior knowledge or retrieval of the published answer. We introduce BioStudyBench, a benchmark of 25 long-horizon analysis tasks drawn from studies first published between July and September 2026, after the developer-reported knowledge cutoffs of the models we evaluate, semi-automatically filtered down from 404,019 PubMed records. In each task, the agent receives a neutral research question but no data files, so it must find and download the relevant public data, search the literature through tools that return only records dated before its cutoff, and report findings through data analysis. To measure gains over prior knowledge, we run every task both with and without access to data and tools. Across eight models, access to data and tools raises the pass rate by 47 percentage points on average over the no-data baseline. Open-weight models across sizes trail closed-weight models, with the best open-weight model passing 81.3% of tasks against 94.7% for the best closed-weight model.
| Comments: | Accepted into AgenticLS (NeurIPS 2026 workshop) |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07614 [cs.AI] |
| (or arXiv:2610.07614v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07614 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: David Li [view email]
[v1]
Tue, 6 Oct 2026 02:01:35 UTC (260 KB)
来源:arXiv:cs.AI · arxiv.org