跳到正文
arXiv:cs.CL· Adri\'an Z\'ame\v{c}n\'ik, Mat\v{e}j Kripner, Martin Kouteck\'y, Martin Balko, Jan Greb\'ik, Pavel Hub\'a\v{c}ek, Robert \v{S}\'amal, V\'aclav Rozho\v{n}·· 3 小时前AI 评分73

多智能体系统 Bolzano 无人工引导求解约 200 个开放数学问题

From Expert-Guided Proof Search to Automated Open-Problem Solving

AI 导读

论文介绍开源多智能体系统 Bolzano,它使用并行证明智能体加验证智能体并维护人类可读的研究状态。专家选题的初步使用产生 8 个经领域专家核验证明的结果;随后在约 3800 个从四组论文提取的开放问题上无人工引导运行,解出约 200 个。其中在 STOC 2026 论文实验中回答了 4 个问题,经论文作者确认。论文已被 NeurIPS 2026 MATH-AI 研讨会接收。

正文

View PDF HTML (experimental)

Abstract:Large language models are increasingly contributing to mathematical research, where progress often depends on efficient proof search, incremental improvements and careful verification. We describe Bolzano, a multi-agent open-source system that uses parallel prover agents with a verifier agent and maintains a human-readable research state. Initial manual use on expert-selected problems yielded 8 results whose proofs were checked by domain experts. Motivated by these case studies, we ran Bolzano without problem-specific human guidance on about 3,800 open problems extracted from four sets of papers, solving about 200 open problems. One experiment used papers accepted to STOC 2026, a top conference in theoretical computer science. There, we answered four questions raised in the papers, as confirmed by their authors.
Comments: Accepted at the 6th Workshop on Mathematical Reasoning and AI (MATH-AI), NeurIPS 2026
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.09769 [cs.AI]
  (or arXiv:2610.09769v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.09769

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Matěj Kripner [view email]
[v1] Wed, 7 Oct 2026 09:52:46 UTC (13 KB)

来源:arXiv:cs.CL · arxiv.org