跳到正文
arXiv:cs.AI· Chien-Ping Lu·· 3 小时前

Agentic AI 的验证与自我改进:基础与边界

Verification and Self-Improvement in Agentic AI: Foundations and Limits

AI 导读

一篇 arXiv 论文用带隐藏终端随机性的有界验证框架,区分了 Agentic AI 靠"搜索更久""获得额外支持"和"修改提议与验证方式"三类改进机制,指出性能分数无法分辨它们。

正文

View PDF HTML (experimental)

Abstract:Agentic AI systems can improve by searching longer, receiving additional support, or modifying how they propose and verify outputs. A performance score does not distinguish these mechanisms. We compare these changes through bounded verification with hidden terminal randomness. A stage specifies admissible transcripts, polynomial bounds, an alternating verification protocol, and a terminal checker. Its native reach uses default support; its closure frontier permits all support already admitted by the interface. Under a uniform pointwise probability gap and task-relative soundness, these are well-defined languages. We prove that independent majority amplification preserves both languages, whereas existential acceptance over random tapes can admit incorrect outputs. Exact verification is the zero-randomness case, with placement and completeness results. The randomized-verifier classes satisfy $\Sigma_k^{\mathrm{P}}\subseteq\Sigma_k^{\mathrm{RV}}\subseteq\Sigma_{k+1}^{\mathrm{P}}$; strict enlargement and depth separation require explicit complexity assumptions, while $\mathrm{BPP}=\mathrm{P}$ yields exact companions with the same frontiers. Representation analysis separates invariant acceptance from core-versus-support labels that can change under refactoring. For recursive self-improvement, uniformly bounded self-modification under a common sound interpreter and fixed verification protocol remains within the same verification class. A separate conditional-error budget controls false selection across adaptively chosen candidates. A quota-enforced XOR-synthesis family separates unbounded ratios of search success from changes in the accepted languages; exact and probabilistic audits check the resulting evidence requirements. The framework ties self-improvement claims to obligations on correctness, admissible evidence, verification resources, and selection error.
Comments: 26 pages, 5 figures. Includes proofs and reproducibility artifacts
Subjects: Artificial Intelligence (cs.AI); Computational Complexity (cs.CC)
Cite as: arXiv:2610.10611 [cs.AI]
  (or arXiv:2610.10611v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.10611

arXiv-issued DOI via DataCite

Submission history

From: Chien-Ping Lu [view email]
[v1] Wed, 7 Oct 2026 06:47:49 UTC (90 KB)

来源:arXiv:cs.AI · arxiv.org