arXiv:cs.CL· Jorma Valjakka, Juhani Kivim\"aki, Juha Myll\"ari, Jukka K. Nurminen·· 6 小时前AI 评分51
幻觉检测基准中的标注问题:一项实证评估
The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation
AI 导读
arXiv 论文(arXiv:2610.08026,已被 NeurIPS 2026 Evaluations & Datasets Track 接收)实证研究幻觉检测基准中的标注标准错位问题:自动标注可能在目标为事实正确性时错误地采用参考答案忠实度标准。
正文
Abstract:In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often benchmarked with open-domain question answering (QA) datasets containing questions and corresponding short reference answers. First, an LLM is used to generate answers to questions within the QA dataset. Then, some automated labeling strategy is used to label these answers as hallucinated or not by comparing them with the reference answers in the dataset. This evaluation setting creates a methodological ambiguity between two criteria: reference faithfulness (whether the answer is fully supported by the reference) and factual correctness (whether the answer is free from contradictions and factually false specific claims). In practice, automated labelers may apply the former criterion even when the intended target is the latter. We study this potential criterion mismatch using 900 human-labeled question-answer pairs spanning three commonly used QA datasets and three generator models, with labels targeting answer-level factual correctness. We evaluate lexical similarity metrics, a reference-entailment NLI baseline, and seven LLM judges under controlled prompt variants as automated labelers. Our experiments reveal substantial disagreement both among automated labeling strategies and between these labels and human annotations. Many strategies also exhibit strong directional error biases, and for most judge-generator pairs, replacing a faithfulness-oriented prompt with a factual-correctness prompt improves agreement with human annotations and reduces false-positive dominance, indicating that automated hallucination labels depend strongly on how the target criterion is specified. Label-source choice should therefore be considered a fundamental part of benchmark design and made explicit, validated, and matched with the benchmark goal.
| Comments: | 27 pages. Accepted at the NeurIPS 2026 Evaluations & Datasets Track. Data: this https URL. Code: this https URL |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.08026 [cs.CL] |
| (or arXiv:2610.08026v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08026 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jorma Valjakka [view email]
[v1]
Tue, 6 Oct 2026 09:19:10 UTC (132 KB)
来源:arXiv:cs.CL · arxiv.org