跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Slim Barkallah, Luke Bailey, Kaiyue Wen, Mohammed Abouzaid, Tengyu Ma·· 15 小时前AI 评分45

Pseudo-Formalization:用伪形式化与 Block Verification 实现自动证明验证

Pseudo-Formalization for Automatic Proof Verification

AI 导读

研究者提出 Pseudo-Formalization(PF)证明格式,将自然语言证明拆解为自包含模块,每块声明前提、结论与自然语言证明,再由 LLM 翻译并逐模块独立验证,该算法称为 Block Verification(BV)。

正文

View PDF HTML (experimental)

Abstract:Reliable verification of proofs remains a bottleneck for training and evaluating AI systems on hard mathematical reasoning. Fully formal proofs, in languages like Lean, are easy to verify because they are unambiguous and modular. Most proofs, particularly those written by AI systems, have neither property, and translating them into formal languages remains challenging in many frontier math settings. We propose Pseudo-Formalization (PF), a proof format that captures the modularity and precision of formal proofs while retaining the flexibility of natural language. A Pseudo-Formal proof is decomposed into self-contained modules, each stating its premises, conclusion, and proof in natural language. To verify the correctness of a regular natural language proof, an LLM translates it to Pseudo-Formal and then verifies each module independently, an algorithm we call Block Verification (BV). We evaluate PF+BV on two benchmarks spanning olympiad and research-level mathematics, where it pareto-dominates LLM-as-judge baselines on error-finding precision and recall. To support future work, we release our research-level proof verification benchmark ArxivMathGradingBench.
Comments: 31 pages, code available at this https URL
Subjects: Logic in Computer Science (cs.LO); Machine Learning (cs.LG)
Cite as: arXiv:2605.20531 [cs.LO]
  (or arXiv:2605.20531v3 [cs.LO] for this version)
  https://doi.org/10.48550/arXiv.2605.20531

arXiv-issued DOI via DataCite

Submission history

From: Luke Bailey [view email]
[v1] Tue, 19 May 2026 22:08:51 UTC (676 KB)
[v2] Wed, 17 Jun 2026 20:49:55 UTC (698 KB)
[v3] Thu, 1 Oct 2026 04:40:14 UTC (770 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org