arXiv:cs.AI· Yuho Lee, Jisu Shin, Nicole Hee-Yeon Kim, Jihwan Bang, Juntae Lee, Kyuwoong Hwang, Fatih Porikli, Hwanjun Song·· 4 小时前AI 评分39
重新思考长视频中的 RAG:检索什么、如何使用?——V-RAGBench 与 CARVE 方法
Rethinking RAG in Long Videos: What to Retrieve and How to Use It?
AI 导读
针对长视频 RAG 现有基准可脱离视频作答、且每查询只用单一模态-粒度配置的问题,研究者提出 V-RAGBench 基准与免训练方法 CARVE。V-RAGBench 由小时级视频上的〈查询、证据块、答案〉三元组构成,每个答案依赖唯一证据块,可解耦评估检索与生成;CARVE 并行运行多配置检索器,用块自适应重排序为每个块选出最优配置并带入生成。
正文
Abstract:Retrieval-augmented generation is extending beyond text to long videos, where query-relevant chunks can be represented across multiple modalities and temporal granularities. Progress in this setting, VideoRAG, is limited by two gaps: existing benchmarks allow queries to be answered without the video, obscuring retrieval errors, and prior methods apply a single modality-granularity configuration per query, ignoring chunk-level variability. We address both by introducing V-RAGBench, a benchmark of $\langle$query, evidence chunk, answer$\rangle$ triplets over hour-scale videos, in which each answer depends on a unique evidence chunk, enabling decoupled evaluation of retrieval and generation, and CARVE, a training-free method that runs parallel retrievers across configurations and uses chunk-adaptive reranking to select a winning configuration for each chunk, which is then carried into generation. On V-RAGBench, CARVE outperforms eight recent VideoRAG baselines on both stages, with the evidence it supplies to the generator interleaving multiple configurations rather than sharing one. Its gains hold across egocentric and third-person long videos and extend to an expanded configuration space.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.13141 [cs.AI] |
| (or arXiv:2606.13141v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2606.13141 arXiv-issued DOI via DataCite |
Submission history
From: Yuho Lee [view email]
[v1]
Thu, 11 Jun 2026 10:05:49 UTC (2,937 KB)
[v2]
Fri, 2 Oct 2026 05:44:07 UTC (2,906 KB)
来源:arXiv:cs.AI · arxiv.org