跳到正文
arXiv:cs.CL· Matthew J. Vowels, Jamie Cummins·· 5 小时前AI 评分49

Kurate:可扩展的科学生产质量分析

Kurate: Scalable Scientific Quality Analysis

AI 导读

研究团队推出 Kurate,利用 LLM 评估已发表研究的质量,结合论文及其试验注册、协议等相关文档,并将每项判断链接到对应原文段落。该系统对 4347 篇论文(其中 3913 篇为随机试验)在统计功效、因果识别、预注册等 8 个维度打分。与专家标注的 60 份临床试验文档相比,协议与结果发表评分点的匹配分别为 221/242 和 294/370,AC1 为 0.94 和 0.81。

正文

View PDF HTML (experimental)

Abstract:Scientific search systems can find papers that are relevant to a question, but they generally do not assess the quality of the evidence that those papers provide. We present Kurate, a system that uses large language models (LLMs) to assess the quality of published studies. Kurate uses both the paper and its related documents (e.g., the study's trial registration and protocol), and links each of its judgments to the passage of text on which that judgment is based. We applied Kurate to a corpus of 4,347 papers (3,913 of which report randomized trials) and scored each paper on 8 dimensions of study design and reporting: specifically, statistical power, causal identification, preregistration, selective reporting, measurement validity, analysis prespecification, reporting transparency, and conflict of interest and funding. Across the corpus, we found that papers most often exhibited issues with statistical power, selective reporting, and analysis prespecification, although average quality differed between clinical areas. When compared against expert annotations of 60 held-out clinical-trial documents, the information Kurate extracted matched the expert label in 221/242 protocol scorepoints and 294/370 results-publication scorepoints, with AC1 0.94 and 0.81, respectively. Using a well-reputed, high quality clinical trial as a worked example, we show how a single paper's overall grade breaks down into separate judgments, with each linked to specific evidence from the trial's registration, protocol, and published report. Together, these results show that large-scale quality assessment of this kind is feasible, and that it can be used to address meta-scientific research questions.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.07306 [cs.CL]
  (or arXiv:2610.07306v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.07306

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Matthew Vowels [view email]
[v1] Mon, 5 Oct 2026 19:44:07 UTC (2,278 KB)

来源:arXiv:cs.CL · arxiv.org