arXiv:cs.AI· Hanchao Zhou, Jialei Li·· 6 小时前AI 评分50
SCOPE 论文提出以语言模型为策略规划器的认证定理证明方法
SCOPE: Certified Theorem Proving with a Language Model as the Policy Planner
AI 导读
arXiv 论文 SCOPE(State-Conditioned Operator Planning and Execution)让语言模型在算子词表上规划、符号引擎执行数值、编译器生成证明,用于 Lean 等证明助手的认证定理证明。
正文
Abstract:In proof assistants such as Lean, a generated proof must pass machine compilation checks, so evaluation needs no human scoring. Direct generation fails on multi-step numeric propositions: a proof is valid only if every content integer is correct, so the pass rate is bounded by the k-th power of the per-integer accuracy. Controlled corruption across 2,617 reference proofs confirms this power law. SCOPE (State-Conditioned Operator Planning and Execution) enforces the natural division of labor: the model plans over an operator vocabulary, a symbolic engine executes the numerics, and a compiler renders the proof. On a 218-problem suite it certifies 191/218 (87.6%) with a 135M backbone; the 7B DeepSeek-Prover-V1.5-RL certifies 18/218 at 27.5 times the tokens and 37.5 times the wall-clock, and DeepSeek-Prover-V2-7B certifies zero on a bidirectional dual suite. Multi-step thinking costs 6.12 discrete decision actions per problem and produces no natural-language thinking text. Replacing the lagged engine state in the decision frame with the current one lifts the pass rate from 117/218 to 191/218, while up-weighting the chain-end loss hurts. On the public Lean-Workbook library, 2,132 of 3,536 gradeable admissible problems certify (60.29%) with zero regression on the main suite. All readings come from a version-frozen review with independent rechecks and reverse verification. Restricting free generation and keeping decision-time information visible is a more direct route than enlarging the model.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08319 [cs.AI] |
| (or arXiv:2610.08319v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08319 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hanchao Zhou [view email]
[v1]
Tue, 6 Oct 2026 13:22:24 UTC (4,931 KB)
来源:arXiv:cs.AI · arxiv.org