arXiv:cs.LG· Hamed Shirzad, Frederik Wenkel, Dominique Beaini, Danica J. Sutherland, Emmanuel Noutahi·· 7 小时前AI 评分32
SeedER:用种子-扩展-检索实现高效知识图谱检索
SeedER: Seed-Expand-Retrieve for Efficient Knowledge Graph Retrieval
AI 导读
研究者提出 SeedER(Seed-Expand-Retrieve),一种基于 Graph Transformer 与强化学习的第一阶段检索器,推理时可在 CPU 上运行,训练时每次只需少量节点的反馈,内存与算力需求远低于用 LLM 智能体或 GNN 处理全图的方法,效果可与微调 LLM 从知识图谱检索信息的方法相比。
正文
Abstract:Knowledge graphs (KGs) offer a rich representation for relational knowledge, but their irregular structure makes retrieval challenging: ego-graph expansion grows rapidly, and dense embedding methods struggle with multi-hop compositional queries. Several approaches use LLM agents to explore the KG, analyze candidate nodes, and decide where to explore next. While expressive, these approaches can incur substantial computational and memory costs. On the other hand, we show theoretically that dense embeddings precomputed for graph nodes, even with augmented structure and neighborhood-aware features, can require embedding dimensions comparable to the size of the graph to answer families of knowledge graph queries. This limitation can be overcome with query-adaptive embeddings under certain conditions, and there are graph neural network (GNN) variants that can do so. However, processing the whole graph with a GNN can also incur substantial memory and computational costs, and it requires dense ground-truth labels indicating whether each node answers the query. Straightforward $k$-hop selection around anchor nodes can also be problematic: small $k$ limits the nodes we can see, and even with small $k$ values such as three and four, the $k$-hop subgraph can grow substantially. In this work, we devise a new approach using Graph Transformers and reinforcement learning that is much less demanding in memory and computation, can run on a CPU at inference time as a first-stage retriever, and requires feedback on only a small number of nodes at a time during training. We show that this method is competitive with methods that fine-tune LLMs to retrieve information from KGs. We call our method SeedER (Seed-Expand-Retrieve), and position it primarily as a first-stage retriever that can run with modest computational resources.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2605.23753 [cs.LG] |
| (or arXiv:2605.23753v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.23753 arXiv-issued DOI via DataCite |
Submission history
From: Frederik Wenkel Ph.D. [view email]
[v1]
Fri, 22 May 2026 15:26:31 UTC (494 KB)
[v2]
Mon, 5 Oct 2026 18:38:15 UTC (1,316 KB)
来源:arXiv:cs.LG · arxiv.org