跳到正文
arXiv:cs.AI· Shuyang Cao, Karthik Radhakrishnan, David Rosenberg, Steven Lu, Pengxiang Cheng, Lu Wang, Shiyue Zhang·· 4 小时前AI 评分45

评估大语言模型的检索鲁棒性

Evaluating the Retrieval Robustness of Large Language Models

AI 导读

研究构建了覆盖五个数据集、1891 个样本的基准,用稀疏和稠密检索器评估 11 个 LLM 在 RAG 场景下的检索鲁棒性,并提出三项对应研究问题的鲁棒性指标。结果显示模型整体鲁棒性较高,但跨任务差异显著,是否采用 RAG 需视具体场景而定。Qwen 和 GPT 模型在关闭推理时鲁棒性明显下降,将检索文档作为工具响应可提升 Claude 模型但会损害 Qwen 和 GPT 模型。

正文

View PDF HTML (experimental)

Abstract:Retrieval-augmented generation (RAG) generally enhances large language models' (LLMs) ability to solve knowledge-intensive tasks. But RAG could also lead to performance degradation due to imperfect retrieval and the model's limited ability to leverage retrieved content. In this work, we evaluate the robustness of LLMs in practical RAG setups (henceforth retrieval robustness). We focus on three research questions: (1) whether RAG is always better than non-RAG; (2) whether more retrieved documents always lead to better performance; and (3) whether document order impacts results. To facilitate this study, we establish a benchmark of 1,891 samples spanning five datasets across three task categories, each with documents retrieved using both sparse and dense retrievers. We introduce three robustness metrics, each corresponding to one research question. Our experiments across 11 LLMs show that models achieve generally high retrieval robustness, but robustness varies substantially across tasks, suggesting that the decision to adopt RAG remains a case-by-case consideration. We further examine four additional prompting strategies that vary how models interact with retrieved documents. We find that Qwen and GPT models suffer notable robustness declines when reasoning is disabled, even on single-hop QA tasks, and that providing retrieved documents as tool responses improves Claude models but hurts Qwen and GPT models, highlighting potential issues of the GPT models regardless of their best overall robustness under vanilla prompting.
Comments: 24 pages
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2505.21870 [cs.CL]
  (or arXiv:2505.21870v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2505.21870

arXiv-issued DOI via DataCite

Submission history

From: Shiyue Zhang [view email]
[v1] Wed, 28 May 2025 01:34:31 UTC (658 KB)
[v2] Fri, 2 Oct 2026 02:13:51 UTC (1,092 KB)

来源:arXiv:cs.AI · arxiv.org