跳到正文
arXiv:cs.CL· Benjamin Wilcox, Dawei Gao, Pradeeban Kathiravelu, Douglas Causey, Kewei Sha, Yunhe Feng·· 3 小时前AI 评分37

ArcticQA:评估 LLM 在北极科学问答中弃答能力的数据集与基准

Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science

AI 导读

研究者发布 ArcticQA 数据集与 ArcticAbstain 基准,用 194 道源自北极一手研究的多选题检验大语言模型能否在无正确答案时弃答。

正文

View PDF HTML (experimental)

Abstract:Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate sensitivity to answer availability. We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence. We further develop ArcticAbstain, a paired benchmark comparing answer-present and answer-absent conditions, with the correct answer replaced by a distractor in the latter and an explicit abstention option in both. We evaluate eight models from the Gemini, Claude, and ChatGPT families at high reasoning effort, with three trials per condition, yielding 9,312 recorded responses. Answer-present abstention rates range from 0.0% to 63.0%, whereas replacing the correct answer increases abstention by 5.05 percentage points on average. These findings highlight substantial baseline differences and the need to evaluate abstention frequency and responsiveness jointly. The dataset and benchmark are available at this https URL.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.09446 [cs.CL]
  (or arXiv:2610.09446v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.09446

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yunhe Feng [view email]
[v1] Wed, 7 Oct 2026 04:58:28 UTC (87 KB)

来源:arXiv:cs.CL · arxiv.org