跳到正文
arXiv:cs.AI· Kriti Faujdar, Smit Kadvani·· 4 小时前AI 评分45

不用 GPU 能做到多好?轻量级幻觉检测在问答、对话与摘要任务上的系统性基准测试

How Far Can You Get Without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation

AI 导读

研究系统评测了 ROUGE-L、语义相似度、BERTScore 和基于 FEVER 训练的 DeBERTa NLI 检测器四种 CPU 可行的轻量级幻觉检测方法,并在 HaluEval 的问答、对话、摘要三项任务上各用 2000 条测试实例评估。

正文

View PDF HTML (experimental)

Abstract:Hallucination detection has become a pressing requirement for trustworthy AI deployment at scale. The most accurate detection methods depend on GPU-intensive inference, proprietary API calls, or white-box access to the generating model, putting them out of reach for resource-constrained researchers and practitioners. We explore a practical alternative: how well can hallucination detection perform using only lightweight, CPU-feasible methods built on public models? We benchmark four such detectors, ROUGE-L, semantic similarity, BERTScore, and a Natural Language Inference (NLI) detector based on a FEVER-trained DeBERTa model, together with a score-level ensemble of similarity and NLI. We evaluate them across all three tasks of the HaluEval benchmark: question answering (QA), dialogue, and summarisation. We calibrate on a held-out validation split, evaluate on 2,000 test instances per task, and report bootstrap confidence intervals. The similarity-NLI ensemble is the most consistent method, but absolute performance is highly task-dependent. It ranks best on QA (F1 = 0.792, AUC-ROC = 0.873) and on dialogue (F1 = 0.694, AUC-ROC = 0.749), where NLI is the strongest standalone method; on summarisation every method performs near chance (AUC-ROC between 0.469 and 0.574). We then ask whether that failure is intrinsic to lightweight detection or an artifact of our single-pass design, and find it is largely the latter. Raising the premise budget from 800 to 1600 characters lifts summarisation AUC-ROC from 0.567 to 0.629, and replacing single-pass scoring with sentence-level chunk aggregation reaches 0.683, still on CPU with the same model, though at roughly twenty times the NLI inference. Summarisation remains by far the hardest task, but our results do not support treating lightweight detection as intrinsically unsuited to it.
Comments: Camera-ready version. Accepted to the Findings track of GroundLM 2026 (EMNLP 2026 workshop). Code: this https URL
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2606.29809 [cs.CL]
  (or arXiv:2606.29809v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2606.29809

arXiv-issued DOI via DataCite

Submission history

From: Kriti Faujdar [view email]
[v1] Mon, 29 Jun 2026 05:43:03 UTC (428 KB)
[v2] Fri, 2 Oct 2026 04:37:27 UTC (429 KB)

来源:arXiv:cs.AI · arxiv.org