arXiv:cs.CL· Adri\'an Z\'ame\v{c}n\'ik, Mat\v{e}j Kripner, Martin Kouteck\'y, Martin Balko, Jan Greb\'ik, Pavel Hub\'a\v{c}ek, Robert \v{S}\'amal, V\'aclav Rozho\v{n}·· 3 小时前AI 评分73
多智能体系统 Bolzano 无人工引导求解约 200 个开放数学问题
From Expert-Guided Proof Search to Automated Open-Problem Solving
AI 导读
论文介绍开源多智能体系统 Bolzano,它使用并行证明智能体加验证智能体并维护人类可读的研究状态。专家选题的初步使用产生 8 个经领域专家核验证明的结果;随后在约 3800 个从四组论文提取的开放问题上无人工引导运行,解出约 200 个。其中在 STOC 2026 论文实验中回答了 4 个问题,经论文作者确认。论文已被 NeurIPS 2026 MATH-AI 研讨会接收。
正文
Abstract:Large language models are increasingly contributing to mathematical research, where progress often depends on efficient proof search, incremental improvements and careful verification. We describe Bolzano, a multi-agent open-source system that uses parallel prover agents with a verifier agent and maintains a human-readable research state. Initial manual use on expert-selected problems yielded 8 results whose proofs were checked by domain experts. Motivated by these case studies, we ran Bolzano without problem-specific human guidance on about 3,800 open problems extracted from four sets of papers, solving about 200 open problems. One experiment used papers accepted to STOC 2026, a top conference in theoretical computer science. There, we answered four questions raised in the papers, as confirmed by their authors.
| Comments: | Accepted at the 6th Workshop on Mathematical Reasoning and AI (MATH-AI), NeurIPS 2026 |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.09769 [cs.AI] |
| (or arXiv:2610.09769v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09769 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Matěj Kripner [view email]
[v1]
Wed, 7 Oct 2026 09:52:46 UTC (13 KB)
来源:arXiv:cs.CL · arxiv.org