arXiv:cs.CL· Arjun Balaji·· 3 小时前AI 评分39
基于弃权的认证:小校准预算下思维链验证器的无分布保证
Certified by Abstention: Distribution-Free Guarantees for Chain-of-Thought Verifiers at Small Calibration Budgets
AI 导读
研究用七个开源模型、五种验证器信号和 37,000 条评分思维链轨迹,考察无分布选择性保证在几十到几百个标注问题的校准预算下对 CoT 验证器的效果。核心发现是"弃权有效性":一份很少触发的证书可能有效,但每次使用都是错的——在已知风险模拟中标准证书最多在 0.3% 校准抽样中失败,但在触发的抽样中失败率高达 69%。
正文
Abstract:Signals that predict whether a chain-of-thought (CoT) trace is correct are compared by AUC, but deploying one requires a threshold with a guarantee. We ask what distribution-free selective guarantees deliver for CoT verifiers at realistic calibration budgets of tens to a few hundred labelled problems, using seven open models, five verifier signals and 37,000 graded traces. The central observation is validity by abstention: an $(\alpha,\delta)$-valid procedure that issues a certificate with probability $P_{\rm fire}$ bounds the failure probability of an issued certificate only by $\delta/P_{\rm fire}$, so a certificate that rarely fires can be valid and wrong every time it is used. In a simulation with known risk the standard certificate fails in at most 0.3% of calibration draws but in up to 69% of those in which it fires. A certification floor and a lattice condition for Benjamini-Hochberg conformal selection explain why certificates abstain at these budgets, and the data bear them out: the standard certificate returns nothing or a large accepted set, and an unreadable residual-stream probe buys two to three times the coverage of the readable signals, an edge a cross-fitted reconstruction cannot recover linearly from the readable features. We then give a floor-started fixed-sequence certificate, valid without monotonicity assumptions, that covers more than the Bonferroni certificate on every model-signal pair and raises coverage at the non-vacuous target $0.75\pi_0$ from 0.05 to 0.16, although the floor keeps absolute coverage small. Finally, a certificate cannot see what matters after deployment: under benchmark shift the error among accepted traces tracks the new task's base error, and under best-of-$n$ selection against the verifier it rises past the target while the empirical failure frequency stays below $\delta$, because abstention absorbs the failures.
| Comments: | 22 pages, 7 figures, 12 tables. Under submission at AISTATS 2027 |
| Subjects: | Machine Learning (stat.ML); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09541 [stat.ML] |
| (or arXiv:2610.09541v1 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09541 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Arjun Balaji [view email]
[v1]
Wed, 7 Oct 2026 06:40:58 UTC (170 KB)
来源:arXiv:cs.CL · arxiv.org