跳到正文
arXiv:cs.AI· Mingda Zhang, Wenjin Liu, Tiesunlong Shen, Zikai Xiao, Zhenghong Lin, Qing Xu, Erik Cambria, Xiaoying Tang, Haoran Luo·· 6 小时前AI 评分43

ScienceClaw:跨自然科学与社会科学评测 AI-for-Science 智能体持续自进化

ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences

AI 导读

ScienceClaw 将 AI-for-Science 智能体的持续自进化形式化为固定参数的程序自进化,统一任务求解、科学验证与程序更新。其配套基准 ScienceClaw-Eval 覆盖 23 个学科,通过顺序任务流与独立重置评估,衡量科学正确性、进化增益、保持能力、跨数据集迁移和进化成本。

正文

View PDF HTML (experimental)

Abstract:Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both the natural and social sciences. We formalize ScienceClaw as fixed-parameter program self-evolution that unifies task solving, scientific verification, and program updates. ScienceClaw-Eval spans 23 disciplines and measures scientific correctness, evolutionary gain, retention, cross-dataset transfer, and evolution cost through sequential streams and independent reset evaluation. Our framework repairs executable workflows through multi-turn interaction, converts re-execution-verified failure--success trajectories into linked Skill and Operator candidates, and retains an update only when source-task replay reproduces the repair and independent scientific tasks improve. Code is available at this https URL.
Comments: 28 pages
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.08691 [cs.AI]
  (or arXiv:2610.08691v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.08691

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Mingda Zhang [view email]
[v1] Tue, 6 Oct 2026 17:08:03 UTC (16,396 KB)

来源:arXiv:cs.AI · arxiv.org