跳到正文
arXiv:cs.LG· Lasse B. Strand, Robert Jakob, Kevin O'Sullivan, Markus Kreft·· 7 小时前AI 评分46

Agentic AutoRAG:用推理驱动智能体优化 RAG 流水线

Agentic AutoRAG: RAG Pipeline Optimization through Reasoning-Driven Agents

AI 导读

Agentic AutoRAG 是一个面向 RAG 多目标超参数优化的 LLM 智能体优化器,通过 Diagnoser 将每个失败问题归因到检索或生成阶段、Proposer 基于模型排名与定价知识库选取下一组配置。

正文

View PDF HTML (experimental)

Abstract:Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many interacting choices, from chunking and embedding model to reranking and generation. Existing optimizers, from greedy search to Bayesian optimization, reduce each trial to an aggregate score and search without modeling why a configuration performed as it did, even though the retrieved chunks already provide evidence about whether each failure occurred during retrieval or after it. We introduce Agentic AutoRAG, an LLM-agent optimizer for multi-objective RAG hyperparameter optimization with retrieval-versus-generation failure attribution. It proposes configurations scored on a frozen exam from the corpus: after each trial a Diagnoser attributes each failed question to retrieval or generation, and a Proposer, grounded in a knowledge base of model rankings and pricing, selects the next configuration, weighing accuracy against cost to trace a Pareto frontier. On three multi-hop QA benchmarks it reaches higher LLM-judge accuracy than every baseline we compare, and within its first 10 trials it matches or beats the statistical baselines' full 30-trial judge accuracy. In its cost-aware mode on a real-world healthcare corpus it reaches a median exam accuracy of 77%, above the strongest baseline's 71.5%, at about 58% of that baseline's cost per query, and it matches that 71.5% at about 22% of the cost.
Comments: Accepted at the Second Workshop for REsearch on Agent Language Models (REALM) at EMNLP 2026 and at the Machine Learning for Systems Workshop at NeurIPS 2026. 9 pages plus references and appendix (16 pages total), 4 figures, 6 tables. Code: this https URL
Subjects: Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)
Cite as: arXiv:2610.08452 [cs.CL]
  (or arXiv:2610.08452v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.08452

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Lasse B. Strand [view email]
[v1] Tue, 6 Oct 2026 14:38:15 UTC (205 KB)

来源:arXiv:cs.LG · arxiv.org