跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Zelin Zhao, Bo Yuan, Yuchen Zhu, Jaemoo Choi, Yongxin Chen·· 15 小时前AI 评分40

RMA:上下文编排的数学研究智能体框架

RMA: Context-Orchestrated Research Math Agents

AI 导读

研究团队提出 Research Math Agents(RMA),一个面向长时程证明开发的智能体框架,核心是持久化类型化研究存储与 Research Context Orchestrator 状态管理层,为每次局部证明操作编译有界上下文。

正文

View PDF

Abstract:Long-horizon mathematical reasoning fails less often because a model cannot produce a valid next step than because an agent fails to maintain and expose the right semantic state across many iterations. Left unmanaged, this produces research-level proofs that are locally convincing yet globally incomplete: a key lemma unproved, an assumption unchecked, a citation unsupported, or a computational claim unverified. We present Research Math Agents (RMA), an agentic framework for long-horizon proof development built around a persistent, typed research store and an orchestrator that compiles operation-specific context from that store. The Research Context Orchestrator is the central state-management layer between the persistent research store and each locally scoped proof operation: it retrieves task-relevant artifacts, compiles them into a bounded context, invokes the appropriate operation, and writes the resulting proof edits, issue updates, literature notes, plans, or evaluations back to the store. This process is designed to keep proof revisions, unresolved issues, prior attempts, literature, and evaluations available across rounds while exposing only task-relevant state to each local operation. We evaluate RMA across complementary research-level settings using independent expert evaluation, blind mathematician review, LLM-based benchmark evaluation, and Lean 4 kernel verification. RMA achieves a 42.5% solve rate on the independently evaluated SOOHAK Challenge Hard set, obtains 8 of 10 correct solutions on First Proof B1 and 8 of 10 passing solutions on B2 under human-expert evaluation, and verifies 213 of 300 sampled Research Solved targets in Formal Conjectures with the Lean 4 kernel.
Comments: Code: this https URL
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2605.22875 [cs.AI]
  (or arXiv:2605.22875v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2605.22875

arXiv-issued DOI via DataCite

Submission history

From: Zelin Zhao [view email]
[v1] Wed, 20 May 2026 04:54:22 UTC (136 KB)
[v2] Wed, 30 Sep 2026 21:05:55 UTC (776 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org