跳到正文
arXiv:cs.LG· Mohamed Aly Bouke·· 5 小时前AI 评分43

句子级上下文敏感性:一种免训练的无支撑内容检测器,与训练型验证器对比评估

Sentence-Level Context Sensitivity as a Training-Free Detector of Unsupported Content, Evaluated Against Trained Verifiers

AI 导读

研究者将"有/无上下文似然对比"这一已知忠实度信号实现为免训练检测器,通过逐一移除检索片段并重打分,定位使句子似然下降最多的候选支撑段落。

正文

View PDF HTML (experimental)

Abstract:Retrieval-augmented generation (RAG) assistants summarize records in clinical and legal work, where one unsupported sentence can mislead a reader. The contrast between an output's likelihood with and without its source is an established faithfulness score for whole summaries and answers, but it has not been measured as a detector of the individual unsupported sentence in multi-passage RAG answers, against trained verifiers, or for its cost. We implement it as a training-free detector that re-scores a fixed answer under the full context, no context, and each chunk removed, and returns the chunk whose removal lowers a sentence's likelihood most as a candidate supporting passage. We evaluate it on RAGTruth, TofuEval, and RAGBench with six scorers and against five verifiers, up to a large language model (LLM) judge, on identical inputs under a source-level split. Scoring per sentence ranks unsupported sentences better than the answer-level form of the same signal on all three benchmarks, by 0.033 to 0.071 in the area under the receiver operating characteristic curve (AUC). On RAGTruth the training-free score reaches an AUC of 0.717 to 0.745 across scorers and 0.773 with a classifier, above entailment and attribution baselines and level with per-chunk fact-checkers, at about one forty-seventh of the LLM judge's compute on a 1.5B scorer, while a full-context fact-checker and the judge are more accurate and are not improved by it. The signal is weakest on short-answer question answering, where the scorer can answer from memory.
Comments: 12 pages. Major revision and retitle of v1 (GASP, arXiv:2607.04223): recast as a controlled evaluation of a known with/without-context likelihood signal; results regenerated under a source-level split with identical inputs; adds an answer-level baseline, a cost analysis, and an annotator study. Code: this https URL
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2607.04223 [cs.CL]
  (or arXiv:2607.04223v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2607.04223

arXiv-issued DOI via DataCite

Submission history

From: Mohamed Aly Bouke [view email]
[v1] Sun, 5 Jul 2026 10:35:30 UTC (970 KB)
[v2] Fri, 2 Oct 2026 14:30:32 UTC (433 KB)

来源:arXiv:cs.LG · arxiv.org