arXiv:cs.LG· Mohamed Aly Bouke·· 5 小时前AI 评分43
句子级上下文敏感性:一种免训练的无支撑内容检测器,与训练型验证器对比评估
Sentence-Level Context Sensitivity as a Training-Free Detector of Unsupported Content, Evaluated Against Trained Verifiers
AI 导读
研究者将"有/无上下文似然对比"这一已知忠实度信号实现为免训练检测器,通过逐一移除检索片段并重打分,定位使句子似然下降最多的候选支撑段落。
正文
Abstract:Retrieval-augmented generation (RAG) assistants summarize records in clinical and legal work, where one unsupported sentence can mislead a reader. The contrast between an output's likelihood with and without its source is an established faithfulness score for whole summaries and answers, but it has not been measured as a detector of the individual unsupported sentence in multi-passage RAG answers, against trained verifiers, or for its cost. We implement it as a training-free detector that re-scores a fixed answer under the full context, no context, and each chunk removed, and returns the chunk whose removal lowers a sentence's likelihood most as a candidate supporting passage. We evaluate it on RAGTruth, TofuEval, and RAGBench with six scorers and against five verifiers, up to a large language model (LLM) judge, on identical inputs under a source-level split. Scoring per sentence ranks unsupported sentences better than the answer-level form of the same signal on all three benchmarks, by 0.033 to 0.071 in the area under the receiver operating characteristic curve (AUC). On RAGTruth the training-free score reaches an AUC of 0.717 to 0.745 across scorers and 0.773 with a classifier, above entailment and attribution baselines and level with per-chunk fact-checkers, at about one forty-seventh of the LLM judge's compute on a 1.5B scorer, while a full-context fact-checker and the judge are more accurate and are not improved by it. The signal is weakest on short-answer question answering, where the scorer can answer from memory.
| Comments: | 12 pages. Major revision and retitle of v1 (GASP, arXiv:2607.04223): recast as a controlled evaluation of a known with/without-context likelihood signal; results regenerated under a source-level split with identical inputs; adds an answer-level baseline, a cost analysis, and an annotator study. Code: this https URL |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.04223 [cs.CL] |
| (or arXiv:2607.04223v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.04223 arXiv-issued DOI via DataCite |
Submission history
From: Mohamed Aly Bouke [view email]
[v1]
Sun, 5 Jul 2026 10:35:30 UTC (970 KB)
[v2]
Fri, 2 Oct 2026 14:30:32 UTC (433 KB)
来源:arXiv:cs.LG · arxiv.org