跳到正文
arXiv:cs.LG· Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Tianshu Fu, Daren Zha, Jun Xiao·· 3 小时前AI 评分37

ROUTEAUDIT:面向预算化多验证器路由的交互感知识别方法

ROUTEAUDIT: Interaction-Aware Identification for Budgeted Multi-Verifier Routing

AI 导读

ROUTEAUDIT 将验证器路由形式化为契约条件下的识别问题,通过契约格、策略无关响应带和请求级边界三个可测量对象,为每次比较返回归因证书。在两个留出原始尾部缓存上,匹配静态 SF+SA 与级联结果相等,将 0.1797 和 0.1250 的表观增益归因于验证器集合优势。

正文

View PDF HTML (experimental)

Abstract:Adaptive multi-verifier systems are commonly compared through endpoint quality-cost gaps, even when the verifier catalog, availability, accounting, information filtration, or scorer changes with the policy. We formulate verifier routing as a contract-conditioned identification problem. The contract records request support, verifier catalog, realized availability, resource accounting, online filtration, and post-trace scoring; a matched route contrast changes only the policy coordinate. ROUTEAUDIT adds three measurable objects to this contract. A contract lattice averages coordinate increments over every admissible bridge order and reports the resulting attribution together with its path sensitivity. A policy-independent response tape identifies paired sequential contrasts when adaptive policies reveal different observations. For incomplete matching, request-level bounds use whichever potential outcome remains observed and give a sharp finite-population interval. The protocol commits paid observations and ledger events before the oracle join and returns an attribution certificate for each comparison. On two held-out raw-tail caches, matched static SF+SA equals the cascade, assigning the apparent gains of 0.1797 and 0.1250 over full static to the verifier-set edge. On 1,319 held-out task requests, the learned and RLVR studies report quality 0.9522 and 0.9553 versus 0.9484 for matched static; the RLVR-static paired difference is +0.0068 with a request-paired interval $[0.0015,0.0122]$ and a training-seed-by-request hierarchical interval $[0.0006,0.0131]$. Controlled attribution recovery yields route mean absolute error 0.0011 and endpoint reconstruction error 0.0004. Factorial, bridge-order, and stochastic-provider studies evaluate the certificate interface; RLVR supplies a learned-policy stress test under the same identification contract.
Comments: 45 pages, 15 figures
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2610.02808 [cs.AI]
  (or arXiv:2610.02808v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.02808

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Miaobo Hu [view email]
[v1] Fri, 2 Oct 2026 04:59:24 UTC (1,533 KB)

来源:arXiv:cs.LG · arxiv.org