arXiv:cs.LG· Jerzy Kami\'nski, Ilya Galyukshev, Artem Kuznetsov, Danil Fedorov, Kirill Redko, Sergey Chuprin, Aidar Shumbalov, Stanislav Chumakov, Anna Kalyuzhnaya·· 7 小时前AI 评分33
大语言模型潜空间"无法回答"信号的跨域与多轮泛化研究
Knowing When Not to Answer: Cross-Domain and Multi-Turn Generalization of Latent Underspecification Signals
AI 导读
研究者构建了一个轮次标注的多轮基准(423 段对话、1,661 个标注轮次状态)和带模拟用户的评测框架,用六个数据集和六个开源权重 LLM 检验"无法回答"探针的泛化能力。
正文
Abstract:Large language models routinely answer questions that cannot be answered from the information given, and in dialogue they answer before enough has been said. Unanswerability is linearly decodable from hidden states, but it is unclear which of its forms share a representation and whether the signal is useful in dialogue. We contribute a turn-labeled multi-turn benchmark (423 conversations, 1,661 labeled turn-states) and an evaluation harness with a simulated user who answers clarifying questions, and use them with six datasets and six open-weight LLMs to test how far probes for unanswerability carry. Probes transfer robustly between datasets that share a ground of unanswerability: missing information in math (AUROC 0.77-0.97) and in a passage (SQuAD 2.0<->MuSiQue, 0.77-0.90). Probes for epistemic "known-unknowns" transfer poorly to math, but this separation weakens under lexical controls and changes with layer and coordinate system, so it remains unresolved. Single-turn probes fail zero-shot to detect when a conversation becomes answerable; in-structure probes recover it, but no better than a bag-of-words classifier. A gate on the calibrated probe, with no model fine-tuning, fires on underspecified turns far more precisely than chance, and its end-task success comes within 0.08 of a gate given the true labels. Yet across four models it does not reliably beat vanilla generation or prompted consolidation. The remaining gap lies mostly in how models use a clarification, not in detection.
| Comments: | 15 pages, 3 figures, 10 tables. Under review |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.08413 [cs.CL] |
| (or arXiv:2610.08413v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08413 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jerzy Kamiński [view email]
[v1]
Tue, 6 Oct 2026 14:18:13 UTC (435 KB)
来源:arXiv:cs.LG · arxiv.org