arXiv:cs.CL· Amanda Myntti, Jenna Kanerva, Veronika Laippala, Filip Ginter·· 3 小时前AI 评分32
嵌入模型中的检索指令为何失效:Your Prompt Should Do More 研究
Your Prompt Should Do More: Effects of Retrieval Instructions in Embedding Models
AI 导读
研究发现嵌入模型在非对称检索任务中即使面对简单任务指令也可能失效,尤其当查询侧存在干扰项时。作者认为这一现象源于当前嵌入模型的训练与评估设置,通过加入查询侧干扰项进行微调可带来显著改善,且对其他任务影响极小。
正文
Abstract:Prompted embedding models have recently received increasing attention, particularly for retrieval, where detailed retrieval instructions are provided as part of the retrieval prompt. Several new datasets and studies have examined this setting, showing that the current embedding models often struggle to follow such instructions reliably. In this paper, we study the mechanism of how instructions actually affect the representations of retrieval queries in asymmetric retrieval tasks. We show that models can fail to follow even simple task instructions when query-side distractors are included in the evaluation. We hypothesize that this behavior is driven by the training setup of current embedding models and their evaluation, and show that fine-tuning with added query-side distractors leads to substantial improvements, with minimal effect on other tasks.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.10508 [cs.CL] |
| (or arXiv:2610.10508v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10508 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Amanda Myntti [view email]
[v1]
Wed, 7 Oct 2026 17:51:46 UTC (1,729 KB)
来源:arXiv:cs.CL · arxiv.org