跳到正文
arXiv:cs.AI· Binjie Guo, Aisheng Mo, Ruitong Li, Xinle Deng·· 4 小时前AI 评分42

EPOCH:通过证据治理搜索实现可靠发现

EPOCH: Reliable Discovery through Evidence-Governed Search

AI 导读

EPOCH 是一种证据治理架构,通过任务契约、类型化记忆、主动证伪、准入检查与独立重放,对候选方案按其支持主张的强度与范围进行评估。它在 AlgoTune 上取得最佳聚合性能,平均归一化得分 0.65,超过最强基线 0.53,并在内部 Math14 套件上取得最高平均分 0.57。在十个发现问题上,EPOCH 产出了改进的可执行构造、优化算法、反例和证明支持的结果。

正文

View PDF HTML (experimental)

Abstract:AI research agents are increasingly used to search over programs, mathematical constructions, and proofs. However, existing systems typically optimize evaluator feedback without adequately governing how that feedback is interpreted, challenged, and reused. As a result, promising but fragile candidates can be promoted as discoveries, while benchmark improvements, finite certificates, and theorem-level claims are too easily conflated. We introduce EPOCH, an evidence-governed architecture designed to close this gap. EPOCH implements an evidence-governed discovery loop by combining explicit task contracts, typed memory, active falsification, admission checks, and independent replay, so that each candidate is evaluated against the strength and scope of the claim it supports. EPOCH achieves state-of-the-art aggregate performance on AlgoTune, substantially exceeding the strongest baseline in mean normalized score (0.65 vs. 0.53), and attains the highest mean score on the internal Math14 suite (0.57). It further shows favorable held-out behavior under official-test replay and leads the descriptive aggregate on AgentHPO. Across ten discovery problems, EPOCH delivers substantial task-specific advances, including improved executable constructions, optimized algorithms, counterexamples, and proof-supported results. These advances demonstrate its ability to convert search into concrete progress across mathematical and computational domains. Together, the results suggest that evidence governance is a necessary step toward AI research agents that produce not only stronger solutions, but also more trustworthy scientific discoveries.
Comments: 49 pages, 16 figures, including supplementary material
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.06986 [cs.AI]
  (or arXiv:2610.06986v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.06986

arXiv-issued DOI via DataCite

Submission history

From: Binjie Guo [view email]
[v1] Sun, 4 Oct 2026 06:51:12 UTC (541 KB)

来源:arXiv:cs.AI · arxiv.org