跳到正文
arXiv:cs.AI· Liulei Zhang, Dejing Zhou, Chuyue Huang, Guanhua Chen, Yutong Yao, Lidia S. Chao, Chi Man Vong, Derek F. Wong·· 4 小时前

自动化研究如何被评估?一篇基准与评估实践综述

How Is Automated Research Evaluated? A Survey of Benchmarks and Evaluation Practices

AI 导读

一篇被 AACL-IJCNLP 2026 Findings 接收的综述系统梳理了自动化研究系统的评估设计,覆盖文献综述、研究构思、可执行工作流、学术写作与交流、自动同行评审及端到端研究六个目标。

正文

View PDF HTML (experimental)

Abstract:Automated research systems support literature synthesis, ideation, experiments, writing, and peer review, but their evaluation is dispersed across tasks, benchmarks, and studies that are difficult to compare directly. We review this literature from the perspective of evaluation design and evidence, covering six targets: literature synthesis, research ideation, executable workflows, scholarly writing and communication, automatic peer review, and end-to-end research. We compare task construction, evidence sources, evaluators, and scoring procedures to explain the capabilities assessed by different designs. Our synthesis highlights three recurring lessons: output checks, process checks, and human studies provide complementary information; evaluator calibration is specific to the property being assessed; and resource budgets and attempt selection are integral to interpreting performance comparisons. We identify diagnostic evaluation designs and documented gaps in supporting evidence, and translate these comparisons into reporting and audit recommendations for specific evaluation settings. The survey helps readers navigate existing evaluations, select appropriate benchmarks, and design subsequent studies.
Comments: Accepted to Findings of AACL-IJCNLP 2026
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11877 [cs.AI]
  (or arXiv:2610.11877v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.11877

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Liulei Zhang [view email]
[v1] Thu, 8 Oct 2026 12:50:23 UTC (132 KB)

来源:arXiv:cs.AI · arxiv.org