跳到正文
arXiv:cs.CL· Linghao Meng, Feng He, Xuan Yang, Junyuan Mao, Pinze Ren, Deqing Mu, Hesen Yang, Qiankun Li·· 3 小时前AI 评分43

PHRBench:LLM 幻觉后推理的行为评估基准

PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLMs

AI 导读

研究者提出 PHRBench,一个覆盖 4 个领域、18 个大语言模型的行为化幻觉后推理(PHR)基准,通过幻觉顺从、幻觉规避与启发式纠正三类行为刻画每条推理轨迹,而非只看最终答案对错。在 4820 个受控实例中,成功纠正幻觉前提并给出正确答案的情况仍相对少见,且与推理轨迹中更频繁的信念更新相关。幻觉提示词本身的属性对成功恢复有显著预测力,一个轻量预测器达到 0.847 的 AUROC。

正文

View PDF HTML (experimental)

Abstract:Hallucinated information can propagate through multi-stage LLM systems and become part of the context for subsequent reasoning. Existing studies of post-hallucination reasoning (PHR) mainly characterize changes in final outcomes and aggregate reasoning dynamics, leaving how models resolve hallucinated premises at the response level insufficiently understood. In this work, we introduce PHRBench, a controlled benchmark for behaviorally structured PHR across four domains and 18 large language models. PHRBench characterizes each reasoning trajectory independently of final-answer correctness through Hallucination Compliance, Hallucination Avoidance, and Heuristic Correction, and defines an insightful trajectory as successful correction that ultimately reaches the correct answer. Across 4820 controlled instances, we find that successful recovery remains relatively rare and is associated with more frequent belief updates along the reasoning trajectory. We further find that properties of the hallucinated prompt contain substantial predictive signal for successful recovery, with a lightweight predictor achieving an AUROC of 0.847. These findings provide a behavioral view of post-hallucination reasoning, characterizing how LLMs resolve erroneous context and when successful recovery is likely to occur.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.10455 [cs.CL]
  (or arXiv:2610.10455v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.10455

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Linghao Meng [view email]
[v1] Wed, 7 Oct 2026 17:25:23 UTC (7,818 KB)

来源:arXiv:cs.CL · arxiv.org