arXiv:cs.AI· Meghanadh Pulivarthi, Swaraj Kumar Biswal, Kushagra Bhushan, Yatin Nandwani, Sachindra Joshi, Dinesh Raghu·· 3 小时前
RIT-RAG:用检索诱导树导航文档语料库
RIT-RAG: Navigating Document Corpora with Retrieval-Induced Trees
AI 导读
RIT-RAG 将内容检索与结构导航结合,离线为每篇文档构建树结构,查询时检索大量 chunk 并用其位置诱导出跨文档的可管理子树,再由 LLM 智能体选择性读取节点并按需重写查询。在金融、科学和客服基准上,其答案准确率超过 vanilla、图结构和智能体基线;在包含 284 万技术文档网页的新基准 EntQABench 上,三个 LLM 的准确率较最强基线提升 6.8 至 11.4 个点。
正文
Abstract:Retrieval-augmented generation (RAG) grounds language models in external corpora. Agentic RAG enables iterative search, yet exposes the model to isolated chunks without document structure, making it difficult to distinguish relevant evidence from chunks that merely resemble the query. Structure-aware methods such as PageIndex navigate document structure but cannot scale to the structures of large corpora, which do not fit in the LLM context. Hence, they first commit to a single document using a document retriever and cannot recover from a wrong choice. We propose RIT-RAG (Retrieval-Induced Tree RAG), which combines content retrieval with structural navigation. Offline, RIT-RAG builds a tree for each document from its table of contents or sitemap. At query time, it retrieves a broad set of chunks and uses their positions to induce manageable sub-trees, potentially across multiple documents. An LLM agent navigates these sub-trees, selectively reads promising nodes, and reformulates queries when needed. Thus, retrieval proposes where to look, while the agent decides what to read. Across financial, scientific, and customer-support benchmarks, RIT-RAG achieves the highest answer accuracy among vanilla, graph-based, and agentic baselines. On EntQABench, our new benchmark of 2.84 million technical-documentation webpages, it improves accuracy by 6.8 to 11.4 points over the strongest baseline across three LLMs.
| Comments: | 24 pages (main text through Limitations ends on page 9, followed by references and appendix), 9 figures, 16 tables |
| Subjects: | Artificial Intelligence (cs.AI); Information Retrieval (cs.IR) |
| Cite as: | arXiv:2610.11370 [cs.AI] |
| (or arXiv:2610.11370v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11370 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Meghanadh Pulivarthi [view email]
[v1]
Thu, 8 Oct 2026 07:03:17 UTC (116 KB)
来源:arXiv:cs.AI · arxiv.org