arXiv:cs.AI· Xilin Dang, Weilin Ruan, Xue Yang, Jinghao Wang, Xiaowei Hu, Jinpeng Li, Pheng-Ann Heng·· 6 小时前AI 评分42
MedZERO:通过受控知识积累实现开放式医学推理的自进化智能体
MedZERO: Self-Evolving Agents for Open-Ended Medical Reasoning Through Controlled Knowledge Accumulation
AI 导读
MedZERO 是一个面向开放式医学推理的自进化框架,由负责生成前沿医学题目选项的 Examiner 与借助外部知识工具进行多轮证据推理的 Reasoner 组成,并采用受控知识积累机制区分临时探索知识与持久精选知识。该框架在 5 个公开医学推理基准上,以 4B 和 8B 规模基座模型评测,一致超越基座模型及此前自进化基线,相对次优自进化基线最高提升 13.7 个平均准确率百分点。
正文
Abstract:Large language models (LLMs) have shown promise in medical question answering and clinical reasoning, yet their improvement remains constrained by static parametric knowledge and costly expert supervision. Self-evolving agents offer a promising alternative by enabling models to improve through iterative task generation and problem-solving. However, most existing self-evolving methods are designed for easily verifiable domains such as mathematics and coding, where solutions can be checked by exact answers or executable programs. Medical reasoning is fundamentally different: it is open-ended, knowledge-intensive, and often only partially verifiable. We present MedZERO, a self-evolving framework for open-ended medical reasoning. MedZERO couples an Examiner that generates frontier medical question-option pairs with a Reasoner that solves them through evidence-grounded multi-turn reasoning with external knowledge tools. To support reliable, continual improvement, MedZERO adopts controlled knowledge accumulation, which maintains temporary exploratory knowledge and curated persistent knowledge in reasoning. We evaluate MedZERO on five public medical reasoning benchmarks using 4B- and 8B-scale base models under open-ended evaluation. Across all settings, MedZERO consistently outperforms the underlying base models and prior self-evolving baselines, achieving up to 13.7 average accuracy-point gains over the next-best self-evolving baseline.
| Comments: | accepted by NIPS 2026 |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08327 [cs.AI] |
| (or arXiv:2610.08327v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08327 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xilin Dang [view email]
[v1]
Tue, 6 Oct 2026 13:25:55 UTC (617 KB)
来源:arXiv:cs.AI · arxiv.org