arXiv:cs.LG· Qi Liu, Fengming Liang, Yiqun Chen, Erhan Zhang, Jiaxin Mao·· 6 小时前AI 评分37
RELER:在嵌入向量空间用强化学习做检索
Learning to Retrieve via Reinforcement Learning in Embedding Space
AI 导读
研究者提出 RELER(REinforcement Learning for Retrieval),一个让现有嵌入模型直接在嵌入向量空间学习检索并对齐任务奖励的强化学习框架,通过从 vMF 分布采样查询和文档嵌入并用 RLOO 更新编码器。
正文
Abstract:Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables existing embedding models to learn to retrieve directly in embedding space and align to task-specific rewards. We train RELER by sampling unit-length query and document embedding actions from von Mises-Fisher (vMF) distributions centered on normalized encoder outputs, scoring the resulting retrieval or downstream outcomes as rewards, and updating the encoder with REINFORCE using a leave-one-out baseline (RLOO). As exploration in the high-dimensional embedding space is prone to sampling noise, we further propose conditional-mean projection (CMP), which projects each sampled embedding onto the low-dimensional subspace spanned by its encoder output and the candidate embeddings it is compared against, reducing noise in the policy gradient while preserving its expectation. We evaluate RELER on BRIGHT, a benchmark with reasoning-intensive queries that remain challenging for existing embedding models. RELER consistently outperforms InfoNCE and LambdaLoss in average nDCG@10 when post-training BGE-M3 and Qwen3-Embedding backbones. We further evaluate downstream utility through retrieval-augmented generation (RAG), where we adapt only the query encoder while keeping the document index and generator fixed. Across seven QA datasets, jointly optimizing retrieval and answer rewards improves both average retrieval performance and answer quality in RAG.
| Subjects: | Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07731 [cs.IR] |
| (or arXiv:2610.07731v1 [cs.IR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07731 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Fengming Liang [view email]
[v1]
Tue, 6 Oct 2026 04:29:29 UTC (202 KB)
来源:arXiv:cs.LG · arxiv.org