跳到正文
arXiv:cs.CL· Mohamed Chenene, Carlos Rosas-Hinostroza, Anastasia Stasenko, Shani Evenstein Sigalov, Pierre-Carl Langlais·· 6 小时前AI 评分47

Wikidata Search Traces:用于训练知识图谱搜索智能体的数据集

Wikidata Search Traces: A Dataset for Training Knowledge Graph Search Agents

AI 导读

研究者构建了 Wikidata Search Traces 数据集,发布 10,235 条单实体与多跳问题的求解轨迹,以及产生这些轨迹的递归语言模型(RLM)框架,该框架让模型批量调用图查询并将结果保存在持久 Python 状态中。

正文

View PDF HTML (experimental)

Abstract:Wikidata is one of the largest open knowledge bases, yet answering a complex question over it still requires a SPARQL query that names the right entities and properties and chains their relations. Language models offer a natural-language alternative but answer largely from memory, which is least reliable for less prominent entities. We study agents that instead answer by exploring the graph, and argue that two obstacles limit them: the lack of training data recording how a solver explores, and interfaces that add large graph results directly to the model's context. We test three hypotheses: that the difficulty of graph search can be controlled through the structure of a question rather than only through obscure entities or wording; that much of the failure on long-horizon search comes from how retrieved evidence is managed rather than from the model itself; and that, in a suitable environment, open-weight models can match commercial closed ones. We construct multi-hop questions on a frozen Wikidata snapshot by replacing named entities with nested conditions, checking after each expansion that the target remains unique and that every new condition is necessary. We release 10,235 solving traces over single-entity and multi-hop questions, together with the recursive language model (RLM) harness that produced them, in which models batch graph calls, keep results in persistent Python state and interpret selected evidence through sub-calls. On 100 questions, the harness improves both models we ran under both interfaces compared with direct tool calling over the same functions: gpt-6-luna rises from 49 to 61 correct answers, doubling its multi-hop accuracy, and Qwen3.8-27B, an open-weight model served on a single GPU, from 60 to 74.
Comments: Technical Report
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.06650 [cs.CL]
  (or arXiv:2610.06650v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.06650

arXiv-issued DOI via DataCite

Submission history

From: Carlos Rosas-Hinostroza [view email]
[v1] Mon, 5 Oct 2026 16:29:22 UTC (295 KB)
[v2] Tue, 6 Oct 2026 16:10:43 UTC (296 KB)

来源:arXiv:cs.CL · arxiv.org