arXiv:cs.AI· Mingda Zhang, Wenjin Liu, Tiesunlong Shen, Zikai Xiao, Zhenghong Lin, Qing Xu, Erik Cambria, Xiaoying Tang, Haoran Luo·· 6 小时前AI 评分43
ScienceClaw:跨自然科学与社会科学评测 AI-for-Science 智能体持续自进化
ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences
AI 导读
ScienceClaw 将 AI-for-Science 智能体的持续自进化形式化为固定参数的程序自进化,统一任务求解、科学验证与程序更新。其配套基准 ScienceClaw-Eval 覆盖 23 个学科,通过顺序任务流与独立重置评估,衡量科学正确性、进化增益、保持能力、跨数据集迁移和进化成本。
正文
Abstract:Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both the natural and social sciences. We formalize ScienceClaw as fixed-parameter program self-evolution that unifies task solving, scientific verification, and program updates. ScienceClaw-Eval spans 23 disciplines and measures scientific correctness, evolutionary gain, retention, cross-dataset transfer, and evolution cost through sequential streams and independent reset evaluation. Our framework repairs executable workflows through multi-turn interaction, converts re-execution-verified failure--success trajectories into linked Skill and Operator candidates, and retains an update only when source-task replay reproduces the repair and independent scientific tasks improve. Code is available at this https URL.
| Comments: | 28 pages |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08691 [cs.AI] |
| (or arXiv:2610.08691v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08691 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mingda Zhang [view email]
[v1]
Tue, 6 Oct 2026 17:08:03 UTC (16,396 KB)
来源:arXiv:cs.AI · arxiv.org