arXiv:cs.AI· Xiaoyang Wang, Tianrui Wang, Christopher C. Yang·· 6 小时前AI 评分37
ProbeGuard:医疗 LLM 共识如何形成?一致同意也可能全错
Unanimously Wrong: Certified Abstention from How Medical LLM Consensus Forms
AI 导读
研究提出认证弃答框架 ProbeGuard,不再只看共识的最终状态,而是依据共识形成过程决定是否弃答。在 MedQA 上,13.4% 的一致投票答案是错的,且任何基于一致性的信号都无法标记;过程信号把区分正确与错误共识的 AUROC 从随机水平提升到 0.696。
正文
Abstract:In clinical practice, agreement among independent experts is treated as evidence of reliability, and multi-round consensus has become a core mechanism of agentic medical question-answering systems. When such a system must decide whether to trust its own answer, the prevailing signal is again agreement, now among the sampled answers. But agreement is a fragile proxy for correctness. A system can be unanimously wrong, returning the same incorrect answer on every sample, and on these questions agreement-based signals carry no information. The cause is that these signals read only the final state of the consensus and discard how it was reached. Agreement that was reached by resolving disagreement with evidence looks identical, at the end, to agreement that was present from the first sample because every sample shares one misconception. ProbeGuard is a certified abstention framework that bases the abstention decision on how the consensus formed. Process features trace agreement trajectories, minority persistence, and retrieval saturation. For unanimous votes, rationale semantic entropy checks whether the reasons behind the vote cohere, and an active probe retrieves counter-evidence and measures whether the consensus survives. A stratified Learn-then-Test calibration then converts these scores into a distribution-free bound on selective risk. We evaluate ProbeGuard on three medical QA benchmarks and a hard-frontier reference, with a published multi-round agentic RAG substrate, against six abstention baselines. On MedQA, 13.4% of unanimous votes are wrong, and no agreement-based signal can flag them. Process signals raise the discrimination of correct from incorrect consensus from chance to 0.696 AUROC. The certified rule answers six in ten unanimous-layer questions at an observed selective risk of 9.0%, and nine in ten once in-domain calibration data accumulate.
| Comments: | Accepted at the GenAI4Health Workshop at NeurIPS 2026 |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07570 [cs.AI] |
| (or arXiv:2610.07570v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07570 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xiaoyang Wang [view email]
[v1]
Tue, 6 Oct 2026 00:59:31 UTC (387 KB)
来源:arXiv:cs.AI · arxiv.org